Skip to content
Development
Skill

/improve-skill-quality

Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement". Use when an evaluation verdict is a regression or underpowered, when a skill regressed after a change, when /evaluate

BOOST
From plugin
managedcode-dotnet-skills
485190 skills17 agents
Install
$ npx -y skills add managedcode/dotnet-skills --skill improve-skill-quality --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/improve-skill-quality

Context preview

The summary Claude sees to decide when to auto-load this skill.

Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement". Use when an evaluation verdict is a regression or underpowered, when a skill regressed after a change, when /evaluate

SKILL.md

improve-skill-quality.SKILL.md
name: improve-skill-quality
description: Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement". Use when an evaluation verdict is a regression or underpowered, when a skill regressed after a change, when /evaluate reports no results, or when deciding whether a weak skill should be strengthened or retired. Do not use for scaffolding a brand-new skill (use create-skill) or a brand-new eval (use create-skill-test).

Improve Skill Quality

Turn a failing or unconvincing evaluation into a targeted fix. The single most common mistake in this repo is rewriting skill prose in response to a verdict whose real cause was the eval, the fixtures, or the harness. Classify first, then fix.

When to Use

  • An evaluation verdict is a regression, underpowered, or "no credible improvement".
  • A skill wins in the isolated arm but not in the plugin arm, or is reported "not activated".
  • `/evaluate` reports "Evaluation ran but produced no results".
  • A skill scores well but costs too much (tokens, turns, wall time, plugin menu budget).
  • Deciding whether to strengthen or retire a persistently weak skill.

When Not to Use

  • Creating a new skill from scratch — use `create-skill`.
  • Creating a new `eval.yaml` from scratch — use `create-skill-test`.
  • Changing the harness itself (`eng/skill-validator`, `eng/vally-adapter`, `evaluation*.yml`).

Inputs

| Input | Required | Description | |-------|----------|-------------| | Verdict evidence | Yes | The `/evaluate` PR comment, or `results.json` from the run artifacts | | Losing trial transcripts | Yes for content fixes | Baseline vs. skilled output plus the judge's stated reason | | Stimulus-vote W/T/L and repeated-run W/T/L | Yes | Separates cross-task evidence from reliability | | Activation status per arm | Yes | Isolated and plugin activation are different failures |

Workflow

Step 1: Get the evidence before forming a hypothesis

Read [InvestigatingResults.md](../../../eng/vally-adapter/InvestigatingResults.md) for how to download artifacts and read `results.json`. Extract, per failing stimulus:

  • authoritative stimulus-vote W/T/L and separate repeated-run W/T/L
  • activation status in the **isolated** and **plugin** arms, separately
  • the judge's verbatim reason on each losing trial
  • whether any trial errored, timed out, or produced empty output

Do not change skill content until you can quote a losing trial and the judge's reason for it. For the other cause classes the evidence is different: harness failures are diagnosed from the job log and the spec, and power problems from the trial record — neither has a losing trial to quote, and demanding one is what sends people rewriting prose instead.

Step 2: Classify the failure

Work down this table and stop at the first row that matches. Rows are ordered by how often the symptom has been misdiagnosed as a skill-content problem — the fixture row is first because a fixture failure also presents as a setup or reliability failure and gets misfiled as one.

| Symptom | Real cause class | Go to | |---------|------------------|-------| | A fixture does not build, is untracked by git, breaks for the wrong reason, or contradicts itself | Fixture | Step 4 | | No `results.json`, "produced no results", or the spec never loaded | Harness / spec-load | Step 3 | | Trials errored, timed out, or returned empty output | Reliability | Step 3 | | Trajectories unmatched, a trial errored, or the summary disagrees — verdict reported inconclusive | Reliability (not power) | Step 3 | | Positive record (e.g. 16W/8T/1L), comparison conclusive, verdict still not a pass | Statistical power | Step 5 | | Skilled arm equals baseline arm by construction | Eval design | Step 6 | | Activated and lost on quality, judge names a concrete defect | Skill content | Step 7 | | Activated in isolation, not in plugin | Activation / routing | Step 8 | | Not activated in either arm | Frontmatter description | Step 8 | | Wins but costs far more than baseline | Scope and cost | Step 7 |

A verdict is only a *measured* result when the comparison was conclusive: `adapt.mjs` requires zero errored trials, zero unmatched trajectories, and an agreeing summary before it will report a pass or a regression. Confirm that before reading a record as a power problem.

Step 3: Rule out harness and reliability causes

See [references/eval-triage.md](references/eval-triage.md) for the full catalogue. The recurring ones:

  • A spec declaring both `config:` and `defaults:` is rejected by vally, the job still exits 0, and

the PR comment blames "transient infrastructure". Merge them into one `defaults:` block.

  • An errored trial is not automatically a fixture problem — judge-side auth and `session.idle`

failures look identical from the verdict and need harness fixes, not SDK pins.

  • `expect_tools: [bash]` on an advisory question forces a restore or build and turns an answer into

a timeout with no quality gain.

  • Genuine code-generation stimuli need roughly 360s; a timeout yields empty output, which fails

every grader and hides the real quality signal.

  • Unmatched trajectories, an errored trial, or a summary that disagrees make the comparison

**inconclusive**: the remaining matched trials are biased, so the record is not a measured null and must not be read as a power or content problem.

Step 4: Verify the fixtures before touching the skill

Run `python eng/eval-quality/check_eval_quality.py` — it blocks eleven defect classes that can cost a real result here. Then confirm by hand:

  • every fixture behaves as its stimulus assumes — a fixture meant to be healthy builds, and one

meant to be broken fails for the exact reason the stimulus is about and no other;

  • every referenced fixture is in the git index (`git ls-files`), not merely on disk — `.gitignore`

has silently swallowed committed coverage fixtures;

  • a fixture never stat
Read more
Ships withmanagedcode-dotnet-skills

Stop explaining .NET to your AI. Start building. We've all been there: asking Claude to use Entity Framework, only to get EF6 patterns in a .NET 8 project. Explaining to Copilot that Blazor Server and Blazor WebAssembly aren't the same thing.

Get the whole plugin

Other skills on managedcode-dotnet-skills.