agents-md-checker
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Waste report for coding-agent sessions in Claude Code, Codex, Gemini CLI, and OpenCode: finds where tokens, money, and time were wasted, and the fix for each kind of waste. Use when the user asks where tokens or money went or why the bill is so high; about habits that burn
$ npx -y skills add RyanAlberts/best-of-Agent-Harnesses --skill session-waste-report --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/session-waste-reportContext preview
The summary Claude sees to decide when to auto-load this skill.
Waste report for coding-agent sessions in Claude Code, Codex, Gemini CLI, and OpenCode: finds where tokens, money, and time were wasted, and the fix for each kind of waste. Use when the user asks where tokens or money went or why the bill is so high; about habits that burn
name: session-waste-report description: >- Waste report for coding-agent sessions in Claude Code, Codex, Gemini CLI, and OpenCode: finds where tokens, money, and time were wasted, and the fix for each kind of waste. Use when the user asks where tokens or money went or why the bill is so high; about habits that burn tokens, such as re-reading a file that has not changed, huge tool outputs, polling, or cache misses from pauses; about the subagent share of cost; or for failure patterns across sessions: tool errors, permission denials, interrupts, corrections, and repeated calls. Runs locally and reads transcripts only; nothing goes over the network. license: MIT metadata: author: "Ryan Alberts" version: "1.0.0" source: "https://github.com/RyanAlberts/best-of-Agent-Harnesses"
A coding agent sends the whole conversation to the model on every call. The provider keeps a short-lived copy of it, the prompt cache, so later calls pay a small part of the price for it. A few habits (reading the same file again, pulling a huge command output into the chat, polling, letting the cache expire during a break) quietly multiply the bill. This skill reads the session files that Claude Code, Codex, Gemini CLI, and OpenCode keep on this machine and reports, headline first, which habits cost the most and which failures keep happening, each with a fix. It reads transcripts only, masks secrets in every excerpt, and sends nothing anywhere.
high.
after breaks, or what subagents (helper conversations the main agent starts for side tasks) cost.
corrections, loops.
payoff before tuning AGENTS.md or settings.
(https://github.com/ccusage/ccusage).
`tool-design-checker`.
`<skill-dir>` means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Keep the quotes in every command: the path can contain spaces.
1. **Run the report** for the window the user asked about, 30 days by default:
python3 "<skill-dir>/scripts/waste.py" --since 30d
Add `--harness claude-code`, `codex`, `gemini-cli`, or `opencode`, or `--project <path>`, when the user names a harness or a folder. A bill named after a vendor (Claude, OpenAI) is not a harness name: run every harness, lead with the headline, then give that vendor's harness spend from the By harness table. `--since` also takes hours (`12h`) and weeks (`2w`). The script only reads; a month of heavy use takes seconds. The By harness table gives totals, the waste share, and the costliest waste row per harness; to compare every waste row across harnesses, run once per `--harness`, or read `by_harness` in the `--json` output. Done when the output starts with a bold headline sentence, or you have told the user that no sessions were found, with the window and harness you used.
2. **Pick the fixes.** Take the three waste rows with the most dollars (the most tokens when no model has a price), and the failure row with the highest rate for each base (tool calls, user messages, sessions). For each, open `references/fixes.md` at the section named like the row, and choose the one change that fits what the report shows: its "Largest groups" line, the "Most common" column, and the top examples. Done when each chosen row has one concrete change: a line for AGENTS.md, a command or flag, or a sibling skill to run.
3. **Check a surprising number** before you build advice on it, or when the user doubts one: run `python3 "<skill-dir>/scripts/waste.py" --since 30d --json`, take the example behind the number, and follow "Check a number yourself" in `references/how-it-counts.md`. Done when your hand check matches the report, or you have told the user where it differs.
4. **Report** in the shape below.
prices. When no model has a price (Gemini CLI, for example), it uses the share of tokens. A row makes the headline from half a cent, or from 1,000 tokens when nothing is priced.
window. Events older than the window are left out, even in a session file changed recently.
between that could change it.
them again.
hour when the session used the 1-hour cache) that had to write the conversation to the cache again.
carrying only; for rebuilds, the tokens written to the cache again. Carrying means the later calls that r
🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly.
Repo: RyanAlberts/best-of-Agent-Harnesses
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot…
Claim checker that audits a coding agent's statements that tests pass or a build is clean…
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up…
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own…
Regression check for coding agents: shows how the agent behaved before and after each harness…
Rule enforcer that finds which written rules in AGENTS.md, CLAUDE.md, and GEMINI.md a coding…