Skip to content

Skills: Testing & Evaluation

A skill is only as good as its evidence. Test four things: discovery (is it found?), selection (is it chosen only when it should be?), execution (does it follow the steps?), and outcome (is the result good?).

Run npm run qa in this guide’s folder (or copy qa/validate.mjs into your skills repo). It checks:

  • SKILL.md exists and has frontmatter
  • name: present, ≤64 chars, ^[a-z0-9]+(-[a-z0-9]+)*$, matches the folder name, and has no reserved words
  • description: present, ≤1024 chars, contains no XML tags
  • the body is under 500 lines
  • every relative link resolves
  • no backslash paths in links
Host How to check
Cursor Type / in chat: the skill is listed. Or see Customize → Skills.
Claude Code Type /: the skill is listed. Or ask “what skills do you have?”
Claude.ai Settings → Capabilities: the skill shows as enabled

If it isn’t listed, check the folder location, that the folder name equals name, the YAML syntax (a stray tab or a : inside an unquoted value breaks it), then restart the host.

Use a fresh chat for each prompt, because earlier context skews selection.

# Prompt Should trigger? Did it? Notes
T1 “Write release notes for v2.4” Yes
T2 “What shipped this sprint? Make it customer friendly” Yes
T3 “Draft the changelog from these PRs” Yes
N1 “Write a commit message for this diff” No near miss
N2 “Summarise this PR for the reviewer” No near miss

Target: 100% on the “Yes” rows and 0 false positives on the “No” rows.

  • If it misses real triggers, add those phrases to the description.
  • If it fires on near misses, narrow the description or state what it’s not for.

For each positive prompt, check:

  • It read SKILL.md, and only the reference files it needed
  • It followed the workflow steps in order
  • It ran scripts rather than rewriting them
  • It completed the validation loop
  • The output matches the template and passes the quality bar
  • It’s better than the baseline you captured in Step 1

Score outcomes against a rubric (for example 0–2 for each criterion) so you can compare versions.

Store your prompts, expected behaviour, and rubric next to the skill (for example evals/evals.md, or JSON if you automate it). Use templates/eval-template.md as a starting point.

Re-run the eval set:

  • after every change to the skill
  • when a new model is released or becomes your default
  • when the host (Cursor or Claude Code) ships a major update

For skills that matter a lot, run the eval prompts headlessly and check the results:

Terminal window
# Claude Code
claude -p "Write release notes for v2.4 from CHANGELOG.md" --output-format json > out.json
# Cursor CLI
agent -p "Write release notes for v2.4 from CHANGELOG.md" --output-format json > out.json

Then assert on the output with a script: required headings present, no banned words, and so on. You can also use a second model call as a grader against your rubric. The Claude Agent SDK and the Cursor SDK both let you script this loop (see SDK agents).

A fast way to improve a skill:

  1. Agent A (the author) helps you write or refine the skill.
  2. Agent B (a fresh session with the skill installed) does real tasks.
  3. You watch B, note where it stumbles, and bring those observations back to A.
  4. Repeat until B handles the eval set cleanly.

This separates writing the skill from using it, so you catch the gaps that only show up when the agent has no prior context.

Symptom Likely cause Fix
Never triggers Vague description, or wrong location Add trigger words; check the path and folder name
Triggers too often Description too broad Narrow the scope; add “Not for X”
Triggers, but ignores steps Body too long, or steps buried Move the workflow to the top; add a checklist
Reads every reference file No “when to read” cues Add conditional pointers
Script errors Missing dependency or wrong working directory Document dependencies; use paths relative to the skill
Works in Claude Code, not Cursor (or the reverse) Platform-specific field, or a location only one host scans Use .claude/skills/ for both; check the matrix