Skip to content
Development
Skill

/create-skill-test

Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and rubrics, sizing an eval for statistical power, or setting up test fixture files. Handles the Vally eval.yaml

BOOST
From plugin
managedcode-dotnet-skills
485190 skills17 agents
Install
$ npx -y skills add managedcode/dotnet-skills --skill create-skill-test --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/create-skill-test

Context preview

The summary Claude sees to decide when to auto-load this skill.

Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and rubrics, sizing an eval for statistical power, or setting up test fixture files. Handles the Vally eval.yaml

SKILL.md

create-skill-test.SKILL.md
name: create-skill-test
description: Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and rubrics, sizing an eval for statistical power, or setting up test fixture files. Handles the Vally eval.yaml schema, fixture organization, and overfitting avoidance. Do not use for running or debugging existing evals (use improve-skill-quality) nor for skills authoring (use create-skill).

Create Skill Test

Scaffold an evaluation spec (`eval.yaml`) for a skill or agent so it conforms to the Vally schema, passes `skill-validator check` and `check_eval_quality.py`, is powerful enough to return a verdict, and does not overfit to the skill's own wording.

When to Use

  • Creating a new `eval.yaml` for a skill or agent
  • Adding stimuli to an existing eval
  • Sizing an eval so the pass gate can actually be reached
  • Setting up or repairing fixture files alongside an eval
  • Reviewing whether rubric items and graders risk overfitting

When Not to Use

  • Diagnosing a failing or regressed eval — use `improve-skill-quality`
  • Modifying the skill-validator or the evaluation workflows
  • Creating or editing `SKILL.md` files — use `create-skill`

Inputs

| Input | Required | Description | |-------|----------|-------------| | Skill or agent name | Yes | Must exist under `plugins/<plugin>/skills/` or `plugins/<plugin>/agents/` | | Plugin name | Yes | e.g. `dotnet-msbuild` | | Skill content | Yes | Read it — you cannot write non-overfitted rubric items without it | | Failure modes to discriminate | Recommended | Each becomes one stimulus |

Workflow

Step 1: Locate the target and the test directory

tests/<plugin>/<skill-name>/eval.yaml          # skills
tests/<plugin>/agent.<agent-name>/eval.yaml    # agents (the agent. prefix disambiguates)

Verify the target exists at `plugins/<plugin>/skills/<skill-name>/SKILL.md` or `plugins/<plugin>/agents/<agent-name>.agent.md`, and read it.

**Agent evals use the native SDK agent lane.** Vally 0.14 cannot register custom agents, so `agent.*` specs do not run through the skill experiment. The evaluation workflow discovers them separately, runs the target agent through `skill-validator evaluate`, and adapts that evidence into the same schema-versioned result and dashboard pipeline. The distinct-stimulus floor applies to both skill and agent evals.

**Be careful with a skill that sets `disable-model-invocation: true`.** The model cannot invoke it, so the skill is absent from the model-facing skilled arm and any direct eval compares two identical arms. Answer-content graders do not create a difference between those arms. The honest coverage for such skills is dependency-level — through the outcome evals of the skills that load them, and through the plugin arm. For example, `filter-syntax` is covered by the filtered-command scenarios in `tests/dotnet-test/run-tests/eval.yaml`.

Step 2: Write the spec skeleton

The spec is Vally format. Every eval in this repo uses `stimuli:` and `graders:`; `scenarios:` and `assertions:` are a pre-Vally format that no longer loads.

name: <skill-name>
description: Evaluates the <plugin>/<skill-name> skill
type: capability
defaults:
  timeout: 5m
  runs: 1
stimuli:
  - name: <what the agent must accomplish>
    prompt: <natural developer request>
    environment:
      files:
        - src: fixtures/<case>/Project.csproj
          dest: Project.csproj
    graders:
      - type: output-matches
        config:
          pattern: (root cause|underlying issue)
      - type: exit-success
      - type: prompt
    rubric:
      - <outcome the agent should have reached>

> **`defaults:` replaces `config:` — it does not join it.** `config` is a deprecated alias for the > same block and vally **throws** on a spec declaring both. Some existing evals still open with > `config:`; when you change settings, replace it with one `defaults:` block. The failure is > invisible otherwise: the job exits 0 with no verdicts and the PR comment > blames "transient infrastructure".

Step 3: Size the eval for power before writing content

The gate gives each distinct stimulus one vote. Repeated runs for one stimulus collapse to one majority-direction vote and remain available as reliability evidence.

1. **Distinct stimuli ≥ 5**, else the verdict is `underpowered` — never a pass, never a regression. 2. **p ≤ 0.05 on an exact one-sided sign test over *discordant* (non-tie) stimulus votes.** Ties are not discarded; they hold the discordant count down.

| discordant stimulus votes | records that pass | p | |---:|---|---:| | ≤ 4 | none | ≥ 0.0625 | | 5–7 | zero losses only (5W/0L) | 0.031 | | 8 | one loss survivable (7W/1L) | 0.035 |

At exactly 5 stimuli, one tie is fatal because it leaves 4 discordant votes. At 6 stimuli one tie is survivable; at 7, up to two are. A loss is not. Five is an **eligibility floor**, not adequate power. For example, 80% power needs 8 discordant votes only for a true 90% conditional win rate; it needs 18 at 80%, 37 at 70%, and 158 at 60%. Size for the effect and tie rate you need to detect.

Use `runs` for reliability, not task breadth. Vally recommends 3 runs in CI and 5–10 nightly for pass rate, pass@k, pass^k, and flakiness. Extra runs never clear the five-stimulus floor.

Do not set `runs` in `dotnet-skills.experiment.yaml`; experiment overrides overwrite every eval's own value rather than defaulting it.

Step 4: Write stimuli

  • **Name** describes *what* is tested, not *how*.
  • **Prompt** is a natural developer request. Never mention the skill, the agent, or its vocabulary —

cued prompts inflate the overfit score and bias the baseline.

  • Each stimulus should discriminate a **different** property of the skill. Five stimuli covering one

property give arithmetic, not evidence.

  • Give every stimulus a stable, unique `name`. Vally pairs comparison trajectories by

`(stim

Read more
Ships withmanagedcode-dotnet-skills

Stop explaining .NET to your AI. Start building. We've all been there: asking Claude to use Entity Framework, only to get EF6 patterns in a .NET 8 project. Explaining to Copilot that Blazor Server and Blazor WebAssembly aren't the same thing.

Get the whole plugin

Other skills on managedcode-dotnet-skills.