eval-audit
**Locale**: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per…
**Locale**: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per SKILL.md's Global UX Rules. Dimension labels: see the canonical table in SKILL.md.
$ npx -y skills add Evol-ai/SkillCompass --agent claude-codeHow it fires
How this command gets triggered: by you, by Claude, or both.
/eval-compareContext preview
What this command does when you run it.
**Locale**: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per SKILL.md's Global UX Rules. Dimension labels: see the canonical table in SKILL.md.
> **Locale**: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per SKILL.md's Global UX Rules. Dimension labels: see the canonical table in SKILL.md.
For each argument:
**Cross-skill check:** If both arguments use `name@version` syntax and the skill names differ, warn the user that they are comparing different skills ({name_a} vs {name_b}) and the result may lack meaningful reference value, then present the choice:
[Continue comparing / Cancel]
If the user chooses Cancel, stop.
For each version, check `.skill-compass/{name}/manifest.json` for cached evaluation results. Use the **Read** tool to load the manifest.
If cached results exist (matching content_hash): use cached scores. If not: run eval-skill flow on the version to generate fresh results.
Generate a side-by-side comparison:
Version Comparison: sql-optimizer | Dimension | v1.0.0 | v1.0.0-evo.2 | Delta | |-----------------|--------|--------------|--------| | D1 Structure | 6 | 7 | ↑ +1 | | D2 Trigger | 3 | 6 | ↑ +3 * | | D3 Security | 2 | 7 | ↑ +5 * | | D4 Functional | 4 | 4 | → 0 | | D5 Comparative | 3 | 3 | → 0 | | D6 Uniqueness | 7 | 7 | → 0 | |-----------------|--------|--------------|--------| | Overall | 38 | 52 | ↑ +14 | | Verdict | FAIL | CAUTION | |
Significance flag (*): delta > 2 points.
Analyze the pattern of changes:
Output assessment as part of the report.
Skip this step if `--internal` or `--ci` is set.
After the report is printed, output a status line then present the following choice:
✓ Comparison complete. Next step: [Improve weaker version (recommended) / Roll back / Done]
Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
Repo: Evol-ai/SkillCompass
**Locale**: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per…
**Locale**: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per…
- **Recommended model: Claude Opus 4.6** (`claude-opus-4-6`). Directed improvement requires understanding complex rubric feedback and generating precise,…
**Locale**: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per…
**Locale**: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per…
**Locale**: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per…