aspnet-core
Build, debug, modernize, or review ASP.NET Core applications with correct hosting,…
Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and rubrics, sizing an eval for statistical power, or setting up test fixture files. Handles the Vally eval.yaml
$ npx -y skills add managedcode/dotnet-skills --skill create-skill-test --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/create-skill-testContext preview
The summary Claude sees to decide when to auto-load this skill.
Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and rubrics, sizing an eval for statistical power, or setting up test fixture files. Handles the Vally eval.yaml
name: create-skill-test description: Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and rubrics, sizing an eval for statistical power, or setting up test fixture files. Handles the Vally eval.yaml schema, fixture organization, and overfitting avoidance. Do not use for running or debugging existing evals (use improve-skill-quality) nor for skills authoring (use create-skill).
Scaffold an evaluation spec (`eval.yaml`) for a skill or agent so it conforms to the Vally schema, passes `skill-validator check` and `check_eval_quality.py`, is powerful enough to return a verdict, and does not overfit to the skill's own wording.
| Input | Required | Description | |-------|----------|-------------| | Skill or agent name | Yes | Must exist under `plugins/<plugin>/skills/` or `plugins/<plugin>/agents/` | | Plugin name | Yes | e.g. `dotnet-msbuild` | | Skill content | Yes | Read it — you cannot write non-overfitted rubric items without it | | Failure modes to discriminate | Recommended | Each becomes one stimulus |
tests/<plugin>/<skill-name>/eval.yaml # skills tests/<plugin>/agent.<agent-name>/eval.yaml # agents (the agent. prefix disambiguates)
Verify the target exists at `plugins/<plugin>/skills/<skill-name>/SKILL.md` or `plugins/<plugin>/agents/<agent-name>.agent.md`, and read it.
**Agent evals use the native SDK agent lane.** Vally 0.14 cannot register custom agents, so `agent.*` specs do not run through the skill experiment. The evaluation workflow discovers them separately, runs the target agent through `skill-validator evaluate`, and adapts that evidence into the same schema-versioned result and dashboard pipeline. The distinct-stimulus floor applies to both skill and agent evals.
**Be careful with a skill that sets `disable-model-invocation: true`.** The model cannot invoke it, so the skill is absent from the model-facing skilled arm and any direct eval compares two identical arms. Answer-content graders do not create a difference between those arms. The honest coverage for such skills is dependency-level — through the outcome evals of the skills that load them, and through the plugin arm. For example, `filter-syntax` is covered by the filtered-command scenarios in `tests/dotnet-test/run-tests/eval.yaml`.
The spec is Vally format. Every eval in this repo uses `stimuli:` and `graders:`; `scenarios:` and `assertions:` are a pre-Vally format that no longer loads.
name: <skill-name>
description: Evaluates the <plugin>/<skill-name> skill
type: capability
defaults:
timeout: 5m
runs: 1
stimuli:
- name: <what the agent must accomplish>
prompt: <natural developer request>
environment:
files:
- src: fixtures/<case>/Project.csproj
dest: Project.csproj
graders:
- type: output-matches
config:
pattern: (root cause|underlying issue)
- type: exit-success
- type: prompt
rubric:
- <outcome the agent should have reached>> **`defaults:` replaces `config:` — it does not join it.** `config` is a deprecated alias for the > same block and vally **throws** on a spec declaring both. Some existing evals still open with > `config:`; when you change settings, replace it with one `defaults:` block. The failure is > invisible otherwise: the job exits 0 with no verdicts and the PR comment > blames "transient infrastructure".
The gate gives each distinct stimulus one vote. Repeated runs for one stimulus collapse to one majority-direction vote and remain available as reliability evidence.
1. **Distinct stimuli ≥ 5**, else the verdict is `underpowered` — never a pass, never a regression. 2. **p ≤ 0.05 on an exact one-sided sign test over *discordant* (non-tie) stimulus votes.** Ties are not discarded; they hold the discordant count down.
| discordant stimulus votes | records that pass | p | |---:|---|---:| | ≤ 4 | none | ≥ 0.0625 | | 5–7 | zero losses only (5W/0L) | 0.031 | | 8 | one loss survivable (7W/1L) | 0.035 |
At exactly 5 stimuli, one tie is fatal because it leaves 4 discordant votes. At 6 stimuli one tie is survivable; at 7, up to two are. A loss is not. Five is an **eligibility floor**, not adequate power. For example, 80% power needs 8 discordant votes only for a true 90% conditional win rate; it needs 18 at 80%, 37 at 70%, and 158 at 60%. Size for the effect and tie rate you need to detect.
Use `runs` for reliability, not task breadth. Vally recommends 3 runs in CI and 5–10 nightly for pass rate, pass@k, pass^k, and flakiness. Extra runs never clear the five-stimulus floor.
Do not set `runs` in `dotnet-skills.experiment.yaml`; experiment overrides overwrite every eval's own value rather than defaulting it.
cued prompts inflate the overfit score and bias the baseline.
property give arithmetic, not evidence.
`(stim
Stop explaining .NET to your AI. Start building. We've all been there: asking Claude to use Entity Framework, only to get EF6 patterns in a .NET 8 project. Explaining to Copilot that Blazor Server and Blazor WebAssembly aren't the same thing.
Repo: managedcode/dotnet-skills
Build, debug, modernize, or review ASP.NET Core applications with correct hosting,…
Build, upgrade, and operate Aspire 13.5.x C# or TypeScript application hosts with the current…
Build, review, or migrate Azure Functions in .NET with correct execution model, isolated…
Build and review Blazor applications across server, WebAssembly, web app, and hybrid…
Maintain or migrate EF6-based applications with realistic guidance on what to keep, what to…
Design, tune, or review EF Core data access with proper modeling, migrations, query…