
Improve Skill Quality
FreeDiagnose and fix underperforming AI skills.
Free · Opens the source repo
What Improve Skill Quality does
Improve Skill Quality is a targeted tool designed for developers working with the dotnet/skills repository. This skill addresses issues where AI skills fail to meet their baseline performance, either by returning inadequate evaluation results or failing to activate properly. It provides a systematic approach to diagnose the underlying causes of skill failures, ensuring that developers can effectively identify whether the issue lies in the skill content, evaluation design, or other factors such as fixture reliability.
The workflow begins by gathering evidence from evaluation results to classify the failure accurately. Developers are guided through a series of steps that help them rule out common issues related to harness and reliability before making any changes to the skill itself. This structured approach minimizes the risk of misdiagnosis, which is often a pitfall when developers attempt to rewrite skill content without fully understanding the root cause of the problem.
This skill is particularly useful when an evaluation verdict shows regression, is underpowered, or when it indicates that no credible improvement has been made. It is also applicable when developers need to decide whether to strengthen or retire a skill that consistently underperforms. By following the outlined steps, users can ensure that they are not only addressing the symptoms of the problem but also the actual causes, leading to more robust and effective AI skills.
Overall, Improve Skill Quality is an essential tool for anyone involved in the development and maintenance of AI skills within the dotnet/skills ecosystem, providing a clear path to diagnose and rectify performance issues efficiently.
When to use it
Use this skill when evaluation results indicate a regression or underperformance in AI skills.
When not to use it
Do not use this skill for creating new skills or evaluations from scratch; it is focused on diagnosing existing issues.
What you can build with it
Diagnosing a Regression Verdict
When an evaluation shows that a skill has regressed, use this tool to analyze the results and identify the underlying issues.
Assessing Activation Failures
If a skill is reported as not activated, this skill helps determine whether the problem lies in the content or the activation process.
Evaluating Skill Performance
When considering whether to strengthen or retire a skill, this tool provides the necessary analysis to make an informed decision.
How to install Improve Skill Quality
View source1. Install with the skills CLI
npx skills add dotnet/skills/improve-skill-quality --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by dotnetImprove Skill Quality
Turn a failing or unconvincing evaluation into a targeted fix. The single most common mistake in this repo is rewriting skill prose in response to a verdict whose real cause was the eval, the fixtures, or the harness. Classify first, then fix.
When to Use
- An evaluation verdict is a regression, underpowered, or "no credible improvement".
- A skill wins in the isolated arm but not in the plugin arm, or is reported "not activated".
/evaluatereports "Evaluation ran but produced no results".- A skill scores well but costs too much (tokens, turns, wall time, plugin menu budget).
- Deciding whether to strengthen or retire a persistently weak skill.
When Not to Use
- Creating a new skill from scratch — use
create-skill. - Creating a new
eval.yamlfrom scratch — usecreate-skill-test. - Changing the harness itself (
eng/skill-validator,eng/vally-adapter,evaluation*.yml).
Inputs
| Input | Required | Description |
|---|---|---|
| Verdict evidence | Yes | The /evaluate PR comment, or results.json from the run artifacts |
| Losing trial transcripts | Yes for content fixes | Baseline vs. skilled output plus the judge's stated reason |
| W/T/L record and trial count | Yes | Distinguishes a real regression from an underpowered eval |
| Activation status per arm | Yes | Isolated and plugin activation are different failures |
Workflow
Step 1: Get the evidence before forming a hypothesis
Read InvestigatingResults.md for how to
download artifacts and read results.json. Extract, per failing stimulus:
- win / tie / loss record and total trials (
trials = stimuli × runs) - activation status in the isolated and plugin arms, separately
- the judge's verbatim reason on each losing trial
- whether any trial errored, timed out, or produced empty output
Do not change skill content until you can quote a losing trial and the judge's reason for it. For the other cause classes the evidence is different: harness failures are diagnosed from the job log and the spec, and power problems from the trial record — neither has a losing trial to quote, and demanding one is what sends people rewriting prose instead.
Step 2: Classify the failure
Work down this table and stop at the first row that matches. Rows are ordered by how often the symptom has been misdiagnosed as a skill-content problem — the fixture row is first because a fixture failure also presents as a setup or reliability failure and gets misfiled as one.
| Symptom | Real cause class | Go to |
|---|---|---|
| A fixture does not build, is untracked by git, breaks for the wrong reason, or contradicts itself | Fixture | Step 4 |
No results.json, "produced no results", or the spec never loaded | Harness / spec-load | Step 3 |
| Trials errored, timed out, or returned empty output | Reliability | Step 3 |
| Trajectories unmatched, a trial errored, or the summary disagrees — verdict reported inconclusive | Reliability (not power) | Step 3 |
| Positive record (e.g. 16W/8T/1L), comparison conclusive, verdict still not a pass | Statistical power | Step 5 |
| Skilled arm equals baseline arm by construction | Eval design | Step 6 |
| Activated and lost on quality, judge names a concrete defect | Skill content | Step 7 |
| Activated in isolation, not in plugin | Activation / routing | Step 8 |
| Not activated in either arm | Frontmatter description | Step 8 |
| Wins but costs far more than baseline | Scope and cost | Step 7 |
A verdict is only a measured result when the comparison was conclusive: adapt.mjs requires zero
errored trials, zero unmatched trajectories, and an agreeing summary before it will report a pass or
a regression. Confirm that before reading a record as a power problem.
Step 3: Rule out harness and reliability causes
See references/eval-triage.md for the full catalogue. The recurring ones:
- A spec declaring both
config:anddefaults:is rejected by vally, the job still exits 0, and the PR comment blames "transient infrastructure". Merge them into onedefaults:block. - An errored trial is not automatically a fixture problem — judge-side auth and
session.idlefailures look identical from the verdict and need harness fixes, not SDK pins. expect_tools: [bash]on an advisory question forces a restore or build and turns an answer into a timeout with no quality gain.- Genuine code-generation stimuli need roughly 360s; a timeout yields empty output, which fails every grader and hides the real quality signal.
- Unmatched trajectories, an errored trial, or a summary that disagrees make the comparison inconclusive: the remaining matched trials are biased, so the record is not a measured null and must not be read as a power or content problem.
Step 4: Verify the fixtures before touching the skill
Run python eng/eval-quality/check_eval_quality.py — it blocks ten defect classes that each already
cost a real result here. Then confirm by hand:
- every fixture behaves as its stimulus assumes — a fixture meant to be healthy builds, and one meant to be broken fails for the exact reason the stimulus is about and no other;
- every referenced fixture is in the git index (
git ls-files), not merely on disk —.gitignorehas silently swallowed committed coverage fixtures; - a fixture never states the same fact in two places that disagree — a Cobertura report whose
declared
line-rate, summary totals and<line>elements differ is the canonical case — or the two arms legitimately read different truths.
Step 5: Check whether the eval could ever have passed
The gate has two independent bars, and confusing them is the usual misdiagnosis:
- Counted trials ≥ 5 (
trials = stimuli × runs). Below that the verdict is reportedunderpowered— never a pass, never a regression. - The sign test must reach p ≤ 0.05 over the discordant (non-tie) trials. Ties are not discarded silently; they hold the discordant count down.
| discordant trials | records that pass | p |
|---|---|---|
| ≤ 4 | none, however good the skill | ≥ 0.0625 |
| 5–7 | zero losses only (5W/0L) | 0.031 |
| 8 | one loss survivable (7W/1L) | 0.035 |
So at exactly 5 counted trials a single tie is fatal — it leaves 4 discordant. At 6 counted trials one tie is survivable (5W/1T/0L); at 7, up to two are (5W/2T/0L). A loss is not.
So a positive record with a failing verdict is a power problem, not a content problem. Fix it by
adding discriminating stimuli (cross-task evidence) rather than raising runs (repetition
only) — except where each stimulus drives an expensive pipeline. Record the reasoning in a comment
above defaults:, as tests/dotnet-test/grade-tests/eval.yaml does.
Step 6: Check whether the two arms differ at all
An eval that compares the skill against itself measures judge noise:
- A dormancy guard (
expect_activation: false) must not also setconstraints.reject_skills. That makes the skilled arm skill-free, i.e. identical to baseline. Across four evals the same guard scored −0.4, +0.4, +0.4 and 0, twice costing a skill its pass. - A skill with
disable-model-invocation: truecannot self-activate, so an eval graded on activation compares two identical arms. Cover it through a consumer skill, or grade the answer content instead, astests/dotnet-test/filter-syntax/eval.yamlandtests/dotnet-test/platform-detection/eval.yamldo. - A grader whose
configis missing its required key enforces nothing, so the stimulus has one fewer assertion than it appears to.
Step 7: Fix skill content against the losing trial
Only now change the skill. Apply the patterns in references/writing-for-baseline-delta.md; the ones that most often flip a loss:
- Replace reference prose the model already knows with decisions it would otherwise get wrong.
- Add stop-conditions so a strong skill does not over-apply — but do not over-correct into answering more narrowly than the baseline did.
- Scale output structure to input size; a dashboard for an 8-test suite loses to a direct answer.
- Require truthful validation reporting; claiming "Build succeeded" after a failed restore is an automatic loss.
- Verify load-bearing API claims by compiling or probing, not by reading source.
- For cost regressions, gate rare or expensive paths behind
references/reads and size any orchestration to the user's scope.
Step 8: Fix activation
Activation failures are frontmatter and routing failures, not body failures. See references/eval-triage.md. Summary:
| Failure | Fix |
|---|---|
| Not activated in any arm | Put the user's own words in description: symptoms, error codes, artifact names, quoted requests |
| A sibling skill wins the prompt | Claim the exact ambiguous words in description, and add matching exclusions on both siblings |
| Model answers with no skill at all | Raise the stakes in the description, de-crowd the plugin menu, verify with the plugin arm |
| Boundary excludes real scenarios | Re-read every "do not use for" clause against every eval prompt and real workflow phase |
| Description at the 1,024-char ceiling | Cut restated body content, not trigger phrases; check the plugin menu budget too |
Step 9: Re-validate
dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- check --plugin ./plugins/<plugin>
python eng/eval-quality/check_eval_quality.py
./eng/run-skill-evals.sh <plugin> <skill>
Then request the official run by submitting a PR review containing /evaluate (Files changed →
Review changes), which binds the run to the reviewed commit. Before declaring a regression on the
result, confirm the skill payload actually changed — reruns on byte-identical content have shifted
7W/2T/2L to 4W/5T/2L.
Validation
- For a content fix, a losing trial and the judge's stated reason are quoted in the PR description.
- The failure was classified before any content was edited.
-
check_eval_quality.pyandskill-validator checkboth pass. - Trial count clears the power bar for the observed tie rate, not just the floor of 5.
- Isolated and plugin activation are both reported.
- The PR body records root cause, fix, and validation so the lesson is reusable.
Common Pitfalls
| Pitfall | Solution |
|---|---|
| Rewriting skill prose in response to an underpowered verdict | Underpowered means too few discordant trials; add discriminating stimuli instead |
Adding defaults: runs: to a spec that already has config: | Merge into a single defaults: block; vally rejects specs with both |
Padding runs to clear the trial floor | Five repeats of one stimulus measure one task; add stimuli |
| Treating an errored trial as fixture nondeterminism | Read the stderr first; judge-side auth failures need harness fixes |
| Fixing a "wrong" answer that the fixture actually made wrong | Check fixture self-consistency before blaming the response |
| Strengthening a skill nobody uses and nothing passes | Weak eval signal plus thin telemetry is a valid retirement case |
| Landing a fix without re-running | Verify the invoked payload contains the fix; judge noise is real |
References
- references/writing-for-baseline-delta.md — content patterns that beat the unskilled model
- references/eval-triage.md — symptom, cause and fix catalogue with PR citations
- eng/eval-quality/README.md — the ten structural gate checks and why each exists
- eng/vally-adapter/InvestigatingResults.md — downloading artifacts and reading
results.json. This is the current guide; the similarly-namedeng/skill-validator/src/docs/InvestigatingResults.mddocuments the retiredskill-validator evaluateschema and does not describe today's results.
Frequently asked questions about Improve Skill Quality
Similar skills
Agent Host Debug Logs
Analyze Agent Host debug logs for deeper insights.
Code OSS Dev - Launch + Debug
Launch and debug Code OSS with isolated profiles.
Phoenix CLI
Debug LLM applications with structured analysis tools.
Power Automate Debugging
Diagnose and fix Power Automate flow errors effectively.
Arize Trace
Inspect and export traces for LLM applications.
Runtime Behavior Probe
Investigate real runtime behavior with precision.
