/evaluation-report
Generate a comprehensive evaluation report for a trained model, highlighting performance, baselines, and subgroup metrics.
$ npx -y skills add yeaight7/agent-powerups --agent claude-codeHow it fires
How this command gets triggered: by you, by Claude, or both.
- Fires itselfClaude auto-loads it when your prompt matches the work.
- You can call itInvoke it directly when you want it.
- Slash command
/evaluation-report
Context preview
What this command does when you run it.
Generate a comprehensive evaluation report for a trained model, highlighting performance, baselines, and subgroup metrics.
Command definition
evaluation-report.mddescription: "Generate a comprehensive evaluation report for a trained model, highlighting performance, baselines, and subgroup metrics." argument-hint: "<model_metrics_json_or_evaluation_log>"
Evaluation Report Command
CRITICAL BEHAVIORAL RULES
1. **Compare to Baseline**: Performance in a vacuum is useless. The report MUST contrast the model's metrics against a naive baseline (e.g., predicting the mean or majority class) or the previous model version. 2. **Highlight the Trade-offs**: Emphasize what the model got worse at (e.g., "Recall improved by 5%, but Precision dropped by 2%").
Execution Steps
1. Parse the `$ARGUMENTS` to extract the metrics. 2. Format a clear Markdown report. 3. Include sections for:
- **Topline Metrics**: The primary KPIs (Accuracy, F1, RMSE, etc.).
- **Baseline Comparison**: How much better is this than a dumb rule?
- **Subgroup/Slice Analysis**: Did performance degrade for specific subsets?
- **Confusion Matrix/Error Analysis**: Where is the model making its most costly mistakes?
4. Output the Markdown report and optionally save it to `reports/YYYY-MM-DD-evaluation.md`.
Curated power-ups for coding agents: skills, slash commands, MCP configs, hooks, AGENTS.md templates, and workflows for serious software engineering. Claude Code, Codex, Antigravity CLI, Cursor and more
Repo: yeaight7/agent-powerups
Other commands on agent-powerups.
- /bug-check
Run automated tests and build checks first, then agent code review. For each bug found, propose or document a regression test.
Open command - /build-fix
Use when a build, type check, or test suite is failing and needs to be unblocked with a minimal change.
Open command - /debug
Use when a bug needs systematic diagnosis before a fix is attempted.
Open command - /doctor
Use to diagnose environment, tooling, and Agent Powerups setup problems.
Open command - /implement
Use to turn a spec or user request into working, tested code.
Open command - /mcp-check
Use to check MCP server prerequisites before activating or using a server.
Open command

