The easiest way to evaluate your Agent Skills. Tests that AI agents correctly discover and use your skills. See examples/ — superlint (simple) and angular-modern (TypeScript grader).
$ npx -y skills add mgechev/skillgrade --agent claude-code
Repo: mgechev/skillgrade
What's inside
The easiest way to evaluate your Agent Skills. Tests that AI agents correctly discover and use your skills.
See examples/ — superlint (simple) and angular-modern (TypeScript grader).

Prerequisites: Node.js 20+, Docker
npm i -g skillgrade
1. Initialize — go to your skill directory (must have SKILL.md) and scaffold:
cd my-skill/
GEMINI_API_KEY=your-key skillgrade init # or ANTHROPIC_API_KEY / OPENAI_API_KEY
# Use --force to overwrite an existing eval.yaml
Generates eval.yaml with AI-powered tasks and graders. Without an API key, creates a well-commented template.
2. Edit — customize eval.yaml for your skill (see eval.yaml Reference).
3. Run:
GEMINI_API_KEY=your-key skillgrade --smoke
The agent is auto-detected from your API key: GEMINI_API_KEY → Gemini, ANTHROPIC_API_KEY → Claude, OPENAI_API_KEY → Codex. Override with --agent=claude.
4. Review:
skillgrade preview # CLI report
skillgrade preview browser # web UI → http://localhost:3847
Reports are saved to $TMPDIR/skillgrade/<skill-name>/results/. Override with --output=DIR.
| Flag | Trials | Use Case |
|---|---|---|
--smoke | 5 | Quick capability check |
--reliable | 15 | Reliable pass rate estimate |
--regression | 30 | High-confidence regression detection |
| Flag | Description |
|---|---|
--eval=NAME[,NAME] | Run specific evals by name (comma-separated) |
--grader=TYPE | Run only graders of a type (deterministic or llm_rubric) |
--trials=N | Override trial count |
--parallel=N | Run trials concurrently |
--agent=gemini|claude|codex|acp|opencode|command | Override agent (default: auto-detect from API key) |
--model=NAME | Model the agent answers with (gemini, claude, codex, opencode). Default: whatever the agent CLI is configured to use |
--provider=docker|local | Override provider |
--acp-command=CMD | ACP agent command (e.g., gemini --acp) |
--command=CMD | Command to run for the command agent (e.g., node mycli.js) |
--opencode-agent=NAME | OpenCode agent (build|plan|explore) |
--opencode-model=MODEL | OpenCode model (provider/model format) |
--output=DIR | Output directory (default: $TMPDIR/skillgrade) |
--validate | Verify graders using reference solutions |
--ci | CI mode: exit non-zero if below threshold |
--threshold=0.8 | Pass rate threshold for CI mode |
--preview | Show CLI results after running |
version: "1"
# Optional: explicit path to skill directory (defaults to auto-detecting SKILL.md)
# skill: path/to/my-skill
defaults:
agent: gemini # gemini | claude | codex | acp | opencode | command
model: claude-opus-5 # model the agent answers with (gemini, claude, codex, opencode)
provider: docker # docker | local
trials: 5
timeout: 300 # seconds
threshold: 0.8 # for --ci mode
grader_model: gemini-3-flash-preview # default LLM grader model
grader_provider: gemini # default LLM grader provider: gemini | anthropic | openai
command: node mycli.js # command to run when agent is 'command' (see Custom Command Agent)
acp: # ACP agent configuration (optional)
command: gemini --acp # command to start ACP-compatible agent
env: # optional environment variables
DEBUG: "1"
docker:
base: node:20-slim
setup: | # extra commands run during image build
apt-get update && apt-get install -y jq
environment: # container resource limits
cpus: 2
memory_mb: 2048
tasks:
- name: fix-linting-errors
instruction: |
Use the superlint tool to fix coding standard violations in app.js.
workspace: # files copied into the container
- src: fixtures/broken-app.js
dest: app.js
- src: bin/superlint
dest: /usr/local/bin/superlint
chmod: "+x"
graders:
- type: deterministic
setup: npm install typescript # grader-specific deps (optional)
run: npx ts-node graders/check.ts
weight: 0.7
- type: llm_rubric
rubric: |
Did the agent follow the check → fix → verify workflow?
provider: gemini # optional: gemini (default) | anthropic | openai
model: gemini-3.5-flash # optional model override
weight: 0.3
# Per-task overrides (optional)
agent: claude
model: claude-sonnet-5 # override the model for this task only
grader_provider: anthropic # override default LLM grader provider
trials: 10
timeout: 600
String values (instruction, rubric, run) support file references — if the value is a valid file path, its contents are read automatically:
instruction: instructions/fix-linting.md
rubric: rubrics/workflow-quality.md
A task can carry the row's answer key and a set of labels:
- name: easy--tooltip-token
instruction: |
Report the background colour token of the Tooltip's default variant.
Write {"token": "..."} to answer.json.
expected: # the answer key — graders only, never the agent
token: Brand/100
variants: [default, hover]
metadata: # labels for --filter, recorded with the results
tier: easy
form: open
tags: [smoke]
graders:
- type: deterministic
run: node graders/check-token.mjs
Both are optional: a task without expected is scored purely on what its
graders measure, so golden-truth and metric-style tasks live in the same suite.
skillgrade delivers expected, it never interprets it — comparison is the
grader's job. A deterministic grader receives the task context as one JSON
document in SKILLGRADE_INPUT, so expected keeps its structure:
// graders/check-token.mjs — one script for every token task
const { task, trial, expected, metadata } = JSON.parse(process.env.SKILLGRADE_INPUT);
const answer = JSON.parse(fs.readFileSync('answer.json', 'utf8')); // cwd is the workspace
console.log(JSON.stringify({
score: answer.token === expected.token ? 1 : 0,
details: `${task}: want ${expected.token}, got ${answer.token ?? '(none)'}`,
}));
# or from a shell grader
want=$(jq -r .expected.token <<< "$SKILLGRADE_INPUT")
An llm_rubric grader gets expected as an ## Expected Output section in its
prompt. Both channels are built after the agent process has exited, and neither
writes to the workspace — unlike run: and rubric:, which are staged into the
agent's working directory before it starts, so keep answer keys out of those.
metadata is what --filter selects on. Values are OR within a key and AND
across keys; --filter and --not-filter are repeatable:
skillgrade --filter=tier=easy,medium # easy OR medium
skillgrade --filter=tier=hard --filter=form=refuse # hard AND refuse
skillgrade --filter=tags=smoke --not-filter=tags=flaky
skillgrade --filter-pattern='^easy--' # regex over task names
skillgrade --filter=tier=easy --list # print the selection, run nothing
A filter on a key that no task declares is an error rather than a silent
match-everything — --filter=teir=easy should not quietly run the whole suite.
Filters work the same whether the tasks are inline or imported.
Any section can live in another file. $import takes a file, a directory (every
.yaml/.yml inside it, sorted), a glob, or a list of those:
version: "1"
defaults:
$import: shared/defaults.yaml # merged in place — keys below win
trials: 3
tasks:
- $import: evals/easy/*.yaml # one task per file, or a file holding a list
- $import: evals/hard # a whole directory
trials: 10 # applied to every task it imports
- name: still-inline # inline tasks keep working
instruction: ...
graders: [...]
Each imported file is a normal YAML document — a single task, a list of tasks, or
an object for a section like defaults. Imported files may import further files;
paths are relative to the file that contains the $import, and cycles are an error.
A task keeps working when you move it into its own file: its relative paths
resolve against its own directory first, then the eval root. So
evals/easy/one.yaml can say instruction: instruction.md for the file next to it
while still pointing run: node graders/check.mjs at the shared graders directory
at the root.
Runs a command and parses JSON from stdout:
- type: deterministic
run: bash graders/check.sh
weight: 0.7
Output format:
{
"score": 0.67,
"details": "2/3 checks passed",
"checks": [
{"name": "file-created", "passed": true, "message": "Output file exists"},
{"name": "content-correct", "passed": false, "message": "Missing expected output"}
]
}
score (0.0–1.0) and details are required. checks is optional.
Bash example:
#!/bin/bash
passed=0; total=2
c1_pass=false c1_msg="File missing"
c2_pass=false c2_msg="Content wrong"
if test -f output.txt; then
passed=$((passed + 1)); c1_pass=true; c1_msg="File exists"
fi
if grep -q "expected" output.txt 2>/dev/null; then
passed=$((passed + 1)); c2_pass=true; c2_msg="Content correct"
fi
score=$(awk "BEGIN {printf \"%.2f\", $passed/$total}")
echo "{\"score\":$score,\"details\":\"$passed/$total passed\",\"checks\":[{\"name\":\"file\",\"passed\":$c1_pass,\"message\":\"$c1_msg\"},{\"name\":\"content\",\"passed\":$c2_pass,\"message\":\"$c2_msg\"}]}"
Use
awkfor arithmetic —bcis not available innode:20-slim.
Evaluates the agent's session transcript against qualitative criteria:
- type: llm_rubric
rubric: |
Workflow Compliance (0-0.5):
- Did the agent follow the mandatory 3-step workflow?
Efficiency (0-0.5):
- Completed in ≤5 commands?
weight: 0.3
provider: gemini # gemini (default) | anthropic | openai
model: gemini-2.0-flash # optional, auto-detected from API key
The provider field selects which LLM API to call:
| Provider | API Key Env Var | Base URL Env Var (optional) | Default Model |
|---|---|---|---|
gemini | GEMINI_API_KEY | - | Dynamically resolved latest Flash model (via API) |
anthropic | ANTHROPIC_API_KEY | ANTHROPIC_BASE_URL | Dynamically resolved latest Haiku model (via API) |
openai | OPENAI_API_KEY | OPENAI_BASE_URL | Dynamically resolved latest Mini/Flash model (via API) |
ANTHROPIC_BASE_URL and OPENAI_BASE_URL enable custom/self-hosted endpoints (Ollama, vLLM, etc.). They apply to both LLM grading and skillgrade init.
graders:
- type: deterministic
run: bash graders/check.sh
weight: 0.7 # 70% — did it work?
- type: llm_rubric
rubric: rubrics/quality.md
weight: 0.3 # 30% — was the approach good?
Final reward = Σ (grader_score × weight) / Σ weight
Use --provider=local in CI — the runner is already an ephemeral sandbox, so Docker adds overhead without benefit.
FAQ
skillgrade is a Claude Code plugin with 2 hand-picked skills for testing work, indexed on Flowy. Install it with the command on its page. It includes skillgrade-graders, skillgrade-setup. Its skills do not fire on their own yet. Request auto-invocation to have Flowy route them as you prompt. Free and open source.
Is this plugin yours?
Claim it with GitHubSubmit a pluginPromote it