bug-check
Run automated tests and build checks first, then agent code review. For each bug found, propose or document a regression test.
Generate a comprehensive evaluation report for a trained model, highlighting performance, baselines, and subgroup metrics.
> /plugin marketplace add yeaight7/agent-powerupsHow it fires
How this command gets triggered: by you, by Claude, or both.
/evaluation-reportContext preview
What this command does when you run it.
Generate a comprehensive evaluation report for a trained model, highlighting performance, baselines, and subgroup metrics.
description: "Generate a comprehensive evaluation report for a trained model, highlighting performance, baselines, and subgroup metrics." argument-hint: "<model_metrics_json_or_evaluation_log>"
1. **Compare to Baseline**: Performance in a vacuum is useless. The report MUST contrast the model's metrics against a naive baseline (e.g., predicting the mean or majority class) or the previous model version. 2. **Highlight the Trade-offs**: Emphasize what the model got worse at (e.g., "Recall improved by 5%, but Precision dropped by 2%").
1. Parse the `$ARGUMENTS` to extract the metrics. 2. Format a clear Markdown report. 3. Include sections for:
4. Output the Markdown report and optionally save it to `reports/YYYY-MM-DD-evaluation.md`.
Curated power-ups for coding agents: skills, slash commands, MCP configs, hooks, AGENTS.md templates, and workflows for serious software engineering. Claude Code, Codex, Antigravity CLI, Cursor and more
Repo: yeaight7/agent-powerups
Run automated tests and build checks first, then agent code review. For each bug found, propose or document a regression test.
Use when a build, type check, or test suite is failing and needs to be unblocked with a minimal change.