aspnet-core
Build, debug, modernize, or review ASP.NET Core applications with correct hosting,…
Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement". Use when an evaluation verdict is a regression or underpowered, when a skill regressed after a change, when /evaluate
$ npx -y skills add managedcode/dotnet-skills --skill improve-skill-quality --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/improve-skill-qualityContext preview
The summary Claude sees to decide when to auto-load this skill.
Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement". Use when an evaluation verdict is a regression or underpowered, when a skill regressed after a change, when /evaluate
name: improve-skill-quality description: Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement". Use when an evaluation verdict is a regression or underpowered, when a skill regressed after a change, when /evaluate reports no results, or when deciding whether a weak skill should be strengthened or retired. Do not use for scaffolding a brand-new skill (use create-skill) or a brand-new eval (use create-skill-test).
Turn a failing or unconvincing evaluation into a targeted fix. The single most common mistake in this repo is rewriting skill prose in response to a verdict whose real cause was the eval, the fixtures, or the harness. Classify first, then fix.
| Input | Required | Description | |-------|----------|-------------| | Verdict evidence | Yes | The `/evaluate` PR comment, or `results.json` from the run artifacts | | Losing trial transcripts | Yes for content fixes | Baseline vs. skilled output plus the judge's stated reason | | Stimulus-vote W/T/L and repeated-run W/T/L | Yes | Separates cross-task evidence from reliability | | Activation status per arm | Yes | Isolated and plugin activation are different failures |
Read [InvestigatingResults.md](../../../eng/vally-adapter/InvestigatingResults.md) for how to download artifacts and read `results.json`. Extract, per failing stimulus:
Do not change skill content until you can quote a losing trial and the judge's reason for it. For the other cause classes the evidence is different: harness failures are diagnosed from the job log and the spec, and power problems from the trial record — neither has a losing trial to quote, and demanding one is what sends people rewriting prose instead.
Work down this table and stop at the first row that matches. Rows are ordered by how often the symptom has been misdiagnosed as a skill-content problem — the fixture row is first because a fixture failure also presents as a setup or reliability failure and gets misfiled as one.
| Symptom | Real cause class | Go to | |---------|------------------|-------| | A fixture does not build, is untracked by git, breaks for the wrong reason, or contradicts itself | Fixture | Step 4 | | No `results.json`, "produced no results", or the spec never loaded | Harness / spec-load | Step 3 | | Trials errored, timed out, or returned empty output | Reliability | Step 3 | | Trajectories unmatched, a trial errored, or the summary disagrees — verdict reported inconclusive | Reliability (not power) | Step 3 | | Positive record (e.g. 16W/8T/1L), comparison conclusive, verdict still not a pass | Statistical power | Step 5 | | Skilled arm equals baseline arm by construction | Eval design | Step 6 | | Activated and lost on quality, judge names a concrete defect | Skill content | Step 7 | | Activated in isolation, not in plugin | Activation / routing | Step 8 | | Not activated in either arm | Frontmatter description | Step 8 | | Wins but costs far more than baseline | Scope and cost | Step 7 |
A verdict is only a *measured* result when the comparison was conclusive: `adapt.mjs` requires zero errored trials, zero unmatched trajectories, and an agreeing summary before it will report a pass or a regression. Confirm that before reading a record as a power problem.
See [references/eval-triage.md](references/eval-triage.md) for the full catalogue. The recurring ones:
the PR comment blames "transient infrastructure". Merge them into one `defaults:` block.
failures look identical from the verdict and need harness fixes, not SDK pins.
a timeout with no quality gain.
every grader and hides the real quality signal.
**inconclusive**: the remaining matched trials are biased, so the record is not a measured null and must not be read as a power or content problem.
Run `python eng/eval-quality/check_eval_quality.py` — it blocks eleven defect classes that can cost a real result here. Then confirm by hand:
meant to be broken fails for the exact reason the stimulus is about and no other;
has silently swallowed committed coverage fixtures;
Stop explaining .NET to your AI. Start building. We've all been there: asking Claude to use Entity Framework, only to get EF6 patterns in a .NET 8 project. Explaining to Copilot that Blazor Server and Blazor WebAssembly aren't the same thing.
Repo: managedcode/dotnet-skills
Build, debug, modernize, or review ASP.NET Core applications with correct hosting,…
Build, upgrade, and operate Aspire 13.5.x C# or TypeScript application hosts with the current…
Build, review, or migrate Azure Functions in .NET with correct execution model, isolated…
Build and review Blazor applications across server, WebAssembly, web app, and hybrid…
Maintain or migrate EF6-based applications with realistic guidance on what to keep, what to…
Design, tune, or review EF Core data access with proper modeling, migrations, query…