Skills: Testing & Evaluation
A skill is only as good as its evidence. Test four things: discovery (is it found?), selection (is it chosen only when it should be?), execution (does it follow the steps?), and outcome (is the result good?).
1. Static checks (automated)
Section titled “1. Static checks (automated)”Run npm run qa in this guide’s folder (or copy qa/validate.mjs into your skills repo). It checks:
SKILL.mdexists and has frontmattername: present, ≤64 chars,^[a-z0-9]+(-[a-z0-9]+)*$, matches the folder name, and has no reserved wordsdescription: present, ≤1024 chars, contains no XML tags- the body is under 500 lines
- every relative link resolves
- no backslash paths in links
2. Discovery test
Section titled “2. Discovery test”| Host | How to check |
|---|---|
| Cursor | Type / in chat: the skill is listed. Or see Customize → Skills. |
| Claude Code | Type /: the skill is listed. Or ask “what skills do you have?” |
| Claude.ai | Settings → Capabilities: the skill shows as enabled |
If it isn’t listed, check the folder location, that the folder name equals name, the YAML syntax (a stray tab or a : inside an unquoted value breaks it), then restart the host.
3. Trigger tests (selection)
Section titled “3. Trigger tests (selection)”Use a fresh chat for each prompt, because earlier context skews selection.
| # | Prompt | Should trigger? | Did it? | Notes |
|---|---|---|---|---|
| T1 | “Write release notes for v2.4” | Yes | ||
| T2 | “What shipped this sprint? Make it customer friendly” | Yes | ||
| T3 | “Draft the changelog from these PRs” | Yes | ||
| N1 | “Write a commit message for this diff” | No | near miss | |
| N2 | “Summarise this PR for the reviewer” | No | near miss |
Target: 100% on the “Yes” rows and 0 false positives on the “No” rows.
- If it misses real triggers, add those phrases to the description.
- If it fires on near misses, narrow the description or state what it’s not for.
4. Execution and outcome tests
Section titled “4. Execution and outcome tests”For each positive prompt, check:
- It read
SKILL.md, and only the reference files it needed - It followed the workflow steps in order
- It ran scripts rather than rewriting them
- It completed the validation loop
- The output matches the template and passes the quality bar
- It’s better than the baseline you captured in Step 1
Score outcomes against a rubric (for example 0–2 for each criterion) so you can compare versions.
5. The eval set
Section titled “5. The eval set”Store your prompts, expected behaviour, and rubric next to the skill (for example evals/evals.md, or JSON if you automate it). Use templates/eval-template.md as a starting point.
Re-run the eval set:
- after every change to the skill
- when a new model is released or becomes your default
- when the host (Cursor or Claude Code) ships a major update
6. Automating evals (optional)
Section titled “6. Automating evals (optional)”For skills that matter a lot, run the eval prompts headlessly and check the results:
# Claude Codeclaude -p "Write release notes for v2.4 from CHANGELOG.md" --output-format json > out.json
# Cursor CLIagent -p "Write release notes for v2.4 from CHANGELOG.md" --output-format json > out.jsonThen assert on the output with a script: required headings present, no banned words, and so on. You can also use a second model call as a grader against your rubric. The Claude Agent SDK and the Cursor SDK both let you script this loop (see SDK agents).
7. Two-agent iteration technique
Section titled “7. Two-agent iteration technique”A fast way to improve a skill:
- Agent A (the author) helps you write or refine the skill.
- Agent B (a fresh session with the skill installed) does real tasks.
- You watch B, note where it stumbles, and bring those observations back to A.
- Repeat until B handles the eval set cleanly.
This separates writing the skill from using it, so you catch the gaps that only show up when the agent has no prior context.
Troubleshooting
Section titled “Troubleshooting”| Symptom | Likely cause | Fix |
|---|---|---|
| Never triggers | Vague description, or wrong location | Add trigger words; check the path and folder name |
| Triggers too often | Description too broad | Narrow the scope; add “Not for X” |
| Triggers, but ignores steps | Body too long, or steps buried | Move the workflow to the top; add a checklist |
| Reads every reference file | No “when to read” cues | Add conditional pointers |
| Script errors | Missing dependency or wrong working directory | Document dependencies; use paths relative to the skill |
| Works in Claude Code, not Cursor (or the reverse) | Platform-specific field, or a location only one host scans | Use .claude/skills/ for both; check the matrix |