/metrics-tokens
Analyze token usage efficiency against the MetaGPT baseline and surface per-step optimization opportunities
$ npx -y skills add jmagly/aiwg --skill metrics-tokens --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/metrics-tokens
Context preview
The summary Claude sees to decide when to auto-load this skill.
Analyze token usage efficiency against the MetaGPT baseline and surface per-step optimization opportunities
SKILL.md
metrics-tokens.SKILL.mdnamespace: aiwg
name: metrics-tokens
platforms: [all]
description: Analyze token usage efficiency against the MetaGPT baseline and surface per-step optimization opportunities
metrics-tokens
You perform deep analysis of token usage efficiency. You compare AIWG workflow token consumption against the MetaGPT 124 tokens/line benchmark (REF-013), identify high-cost operations, and surface optimization opportunities.
Triggers
Alternate expressions and non-obvious activations (primary phrases are matched automatically from the skill description):
- "how efficient are my tokens" → efficiency ratio vs MetaGPT baseline
- "am I above the baseline" → threshold status check
- "where are tokens being wasted" → per-step breakdown with recommendations
- "token ratio" → tokens/line ratio calculation
Trigger Patterns Reference
| Pattern | Example | Action | |---------|---------|--------| | Efficiency report | "token efficiency" | `aiwg metrics-tokens` | | Session analysis | "analyze tokens for this session" | `aiwg metrics-tokens --session current` | | Threshold check | "are we at green" | `aiwg metrics-tokens --threshold` | | Per-step breakdown | "which step used the most tokens" | `aiwg metrics-tokens --by-step` | | Optimization hints | "suggest token optimizations" | `aiwg metrics-tokens --optimize` |
Behavior
When triggered:
1. **Determine scope**:
- Default: current or most recent session
- `--session <name>`: named session
- `--all`: aggregate across all sessions
2. **Load token data**:
- Read `.aiwg/ralph/sessions/*/metrics.json` for raw token counts
- Apply estimation heuristic: 4 chars per token (aligned with `src/metrics/token-counter.ts`)
3. **Compute efficiency metrics**:
- Tokens/line ratio for session output
- `vsBenchmark`: percentage vs MetaGPT 124 tokens/line (negative = better)
- `vsBaseline`: percentage vs typical LLM 200 tokens/line (negative = better)
- Threshold status: green (≤124), yellow (125–150), red (>150)
4. **Run the command**:
# Default efficiency report
aiwg metrics-tokens
# Current session
aiwg metrics-tokens --session current
# Per-step breakdown
aiwg metrics-tokens --by-step
# With optimization suggestions
aiwg metrics-tokens --optimize
# JSON output
aiwg metrics-tokens --json
Benchmark Reference
The MetaGPT 124 tokens/line benchmark comes from REF-013 (research corpus). It represents a validated efficiency target for AI-assisted software workflows. AIWG tracks against this benchmark to make token costs legible and comparable across sessions.
| Threshold | Tokens/Line | Status | Action | |-----------|-------------|--------|--------| | At or below benchmark | ≤ 124 | green | No action needed | | Above benchmark | 125–150 | yellow | Flag for review | | Well above benchmark | > 150 | red | Generate optimization recommendations |
Comparison points:
| Baseline | Tokens/Line | |----------|-------------| | MetaGPT benchmark (REF-013) | 124 | | Typical LLM baseline | ~200 | | AIWG target | ≤ 124 |
Report Format
Standard Efficiency Report
Token Efficiency — Session: sdlc-review-20260401-143022
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Token Counts
Input: 42,310 tokens
Output: 18,940 tokens
Total: 61,250 tokens
Content Metrics
Characters: 245,000
Non-blank lines: 548
Total lines: 621
Efficiency
Tokens/line: 112
vs MetaGPT: -9.7% (better than 124 tokens/line benchmark)
vs LLM baseline: -44% (well below 200 tokens/line typical)
Status: green
Threshold: green — at or below MetaGPT benchmark
Per-Step Breakdown (`--by-step`)
Token Efficiency by Step
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Step Tokens Lines Tokens/Line Status
────────────────────── ──────── ───── ─────────── ──────
architecture-designer 18,200 168 108 green
security-architect 14,600 132 111 green
test-architect 13,100 119 110 green
technical-writer 15,350 129 119 green ← highest volume
──────────────────────────────────
Total 61,250 548 112 greenOptimization Report (`--optimize`)
Optimization Suggestions
━━━━━━━━━━━━━━━━━━━━━━━━
Status: green — no critical optimizations needed.
Opportunities (optional):
1. technical-writer (119 tok/line) — near benchmark ceiling.
Consider: scope the synthesis prompt to final merge only,
avoid re-reading full drafts.
2. architecture-designer (18,200 tokens) — highest absolute cost.
Consider: pass only the relevant SAD section, not the full doc.Efficiency Calculation
Token efficiency uses the estimation and comparison logic from `src/metrics/token-counter.ts`:
tokens = ceil(characters / 4)
tokensPerLine = tokens / nonBlankLines
vsBenchmark = (tokensPerLine - 124) / 124 * 100 (negative = better)
vsBaseline = (tokensPerLine - 200) / 200 * 100 (negative = better)
Examples
Example 1: Quick efficiency check
**User**: "Token efficiency for this session"
**Action**:
aiwg metrics-tokens
**Response**: Efficiency report with tokens/line ratio, benchmark comparison, and green/yellow/red status.
Example 2: Identify expensive steps
**User**: "Which step used the most tokens?"
**Action**:
aiwg metrics-tokens --by-step
**Response**: Per-step table showing token counts, line counts, tokens/line ratio, and threshold status for each workflow step.
Example 3: Optimization pass
**User**: "Suggest ways to reduce token usage"
**Action**:
aiwg metrics-tokens --optimize
**Response**: Optimization suggestions targeted at steps above the green threshold, with specific prompt-scoping recommendations.
Example 4: Are we at green?
**User**: "Are we at green on token efficiency?"
**Ext
Read more
namespace: aiwg name: metrics-tokens platforms: [all] description: Analyze token usage efficiency against the MetaGPT baseline and surface per-step optimization opportunities
metrics-tokens
You perform deep analysis of token usage efficiency. You compare AIWG workflow token consumption against the MetaGPT 124 tokens/line benchmark (REF-013), identify high-cost operations, and surface optimization opportunities.
Triggers
Alternate expressions and non-obvious activations (primary phrases are matched automatically from the skill description):
- "how efficient are my tokens" → efficiency ratio vs MetaGPT baseline
- "am I above the baseline" → threshold status check
- "where are tokens being wasted" → per-step breakdown with recommendations
- "token ratio" → tokens/line ratio calculation
Trigger Patterns Reference
| Pattern | Example | Action | |---------|---------|--------| | Efficiency report | "token efficiency" | `aiwg metrics-tokens` | | Session analysis | "analyze tokens for this session" | `aiwg metrics-tokens --session current` | | Threshold check | "are we at green" | `aiwg metrics-tokens --threshold` | | Per-step breakdown | "which step used the most tokens" | `aiwg metrics-tokens --by-step` | | Optimization hints | "suggest token optimizations" | `aiwg metrics-tokens --optimize` |
Behavior
When triggered:
1. **Determine scope**:
- Default: current or most recent session
- `--session <name>`: named session
- `--all`: aggregate across all sessions
2. **Load token data**:
- Read `.aiwg/ralph/sessions/*/metrics.json` for raw token counts
- Apply estimation heuristic: 4 chars per token (aligned with `src/metrics/token-counter.ts`)
3. **Compute efficiency metrics**:
- Tokens/line ratio for session output
- `vsBenchmark`: percentage vs MetaGPT 124 tokens/line (negative = better)
- `vsBaseline`: percentage vs typical LLM 200 tokens/line (negative = better)
- Threshold status: green (≤124), yellow (125–150), red (>150)
4. **Run the command**:
# Default efficiency report aiwg metrics-tokens # Current session aiwg metrics-tokens --session current # Per-step breakdown aiwg metrics-tokens --by-step # With optimization suggestions aiwg metrics-tokens --optimize # JSON output aiwg metrics-tokens --json
Benchmark Reference
The MetaGPT 124 tokens/line benchmark comes from REF-013 (research corpus). It represents a validated efficiency target for AI-assisted software workflows. AIWG tracks against this benchmark to make token costs legible and comparable across sessions.
| Threshold | Tokens/Line | Status | Action | |-----------|-------------|--------|--------| | At or below benchmark | ≤ 124 | green | No action needed | | Above benchmark | 125–150 | yellow | Flag for review | | Well above benchmark | > 150 | red | Generate optimization recommendations |
Comparison points:
| Baseline | Tokens/Line | |----------|-------------| | MetaGPT benchmark (REF-013) | 124 | | Typical LLM baseline | ~200 | | AIWG target | ≤ 124 |
Report Format
Standard Efficiency Report
Token Efficiency — Session: sdlc-review-20260401-143022 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Token Counts Input: 42,310 tokens Output: 18,940 tokens Total: 61,250 tokens Content Metrics Characters: 245,000 Non-blank lines: 548 Total lines: 621 Efficiency Tokens/line: 112 vs MetaGPT: -9.7% (better than 124 tokens/line benchmark) vs LLM baseline: -44% (well below 200 tokens/line typical) Status: green Threshold: green — at or below MetaGPT benchmark
Per-Step Breakdown (`--by-step`)
Token Efficiency by Step
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Step Tokens Lines Tokens/Line Status
────────────────────── ──────── ───── ─────────── ──────
architecture-designer 18,200 168 108 green
security-architect 14,600 132 111 green
test-architect 13,100 119 110 green
technical-writer 15,350 129 119 green ← highest volume
──────────────────────────────────
Total 61,250 548 112 greenOptimization Report (`--optimize`)
Optimization Suggestions
━━━━━━━━━━━━━━━━━━━━━━━━
Status: green — no critical optimizations needed.
Opportunities (optional):
1. technical-writer (119 tok/line) — near benchmark ceiling.
Consider: scope the synthesis prompt to final merge only,
avoid re-reading full drafts.
2. architecture-designer (18,200 tokens) — highest absolute cost.
Consider: pass only the relevant SAD section, not the full doc.Efficiency Calculation
Token efficiency uses the estimation and comparison logic from `src/metrics/token-counter.ts`:
tokens = ceil(characters / 4) tokensPerLine = tokens / nonBlankLines vsBenchmark = (tokensPerLine - 124) / 124 * 100 (negative = better) vsBaseline = (tokensPerLine - 200) / 200 * 100 (negative = better)
Examples
Example 1: Quick efficiency check
**User**: "Token efficiency for this session"
**Action**:
aiwg metrics-tokens
**Response**: Efficiency report with tokens/line ratio, benchmark comparison, and green/yellow/red status.
Example 2: Identify expensive steps
**User**: "Which step used the most tokens?"
**Action**:
aiwg metrics-tokens --by-step
**Response**: Per-step table showing token counts, line counts, tokens/line ratio, and threshold status for each workflow step.
Example 3: Optimization pass
**User**: "Suggest ways to reduce token usage"
**Action**:
aiwg metrics-tokens --optimize
**Response**: Optimization suggestions targeted at steps above the green threshold, with specific prompt-scoping recommendations.
Example 4: Are we at green?
**User**: "Are we at green on token efficiency?"
**Ext
Multi-agent AI framework for Claude Code, Copilot, Cursor, Warp, and 6 more platforms 200+ agents, 109+ CLI commands, 400+ deployable agent/skill/command/rule artifacts, 8 core frameworks, 32 addons, and a 40-plugin Claude Code marketplace.
Repo: jmagly/aiwg
Other skills on aiwg.
- /agent-loop-ext
Crash-resilient external agent loop with state persistence and CI/CD integration
Open skill - /agent-loop
Detect requests for iterative autonomous agent loops and route to the appropriate loop executor
Open skill - /auto-test-execution
Automatically execute tests when code-generating agents modify source files, enforcing the execute-before-return pattern
Open skill - /cross-task-learner
Enable agent loops to learn from similar past tasks and share patterns across loops
Open skill - /debug-memory
Query and manage the executable feedback debug memory
Open skill - /execute-feedback
Execute tests on generated code and iterate until passing
Open skill

