Skip to content

Step-by-Step: Build a Skill

Follow the steps in order. Each step has a Goal, what to Do, and how to Verify it. Tick off the matching items in the skills checklist as you go.

Running example: a writing-release-notes skill that turns merged PRs into customer-facing release notes.


  • Goal: Know exactly what the skill should achieve.
  • Do:
    1. Write a problem statement: “When someone asks for release notes, the agent produces inconsistent, developer-speak notes. We want customer-facing notes grouped into New / Improved / Fixed, in our voice.”
    2. Collect 3–5 real requests, for example “write release notes for v2.4”, “summarise what shipped this sprint for customers”, or “draft the changelog from these PRs”.
    3. Run those requests without the skill and save the outputs. This is your baseline.
    4. List what the agent got wrong. That list is exactly what your skill needs to contain, and nothing more.
  • Verify: You have a problem statement, the example requests, the baseline outputs, and a list of gaps.

Step 2: Confirm a skill is the right mechanism

Section titled “Step 2: Confirm a skill is the right mechanism”
  • Goal: Avoid building the wrong thing.
  • Do: Walk through the decision tree. A skill is right when the agent can do the task but lacks your procedure or knowledge, and the task comes up sometimes rather than on every request.
  • Verify: You can say in one sentence why this isn’t a rule, a subagent, or an MCP server.
  • Goal: A description the agent reliably matches.
  • Do:
    1. Name: lowercase, hyphens, ≤64 characters, descriptive, and ideally a gerund or noun phrase (writing-release-notes, pdf-form-filling). Avoid helper, utils, and names containing claude or anthropic.

    2. Description: third person, ≤1024 characters, saying what it does and when to use it, with the trigger words users actually type:

      description: >-
      Writes customer-facing release notes from merged PRs, commits, or a
      changelog, grouped into New, Improved and Fixed in the company voice.
      Use when the user asks for release notes, a changelog, "what shipped",
      or a customer update about a version or sprint.
    3. Decide on invocation. Can it run automatically, or should it be disable-model-invocation: true (manual only)?

  • Verify: Read the description in isolation. Would the agent pick it for all of your Step 1 requests, and not for a nearby request such as “write a commit message”?
  • Goal: The right audience can see it.

  • Do: Pick from the platform matrix:

    Audience Location
    Just you, every project, Claude Code + Cursor ~/.claude/skills/<name>/
    Just you, Cursor only ~/.cursor/skills/<name>/
    Everyone in this repo, Claude Code + Cursor .claude/skills/<name>/ (committed)
    Everyone in this repo, vendor-neutral .agents/skills/<name>/ (Cursor, Codex; check your other tools)
    Many repos or teams A plugin (see Distribution)
    Claude.ai / Desktop users ZIP upload
  • Verify: The folder name exactly matches name.

  • Goal: A valid, detectable skill.
  • Do:
    1. Copy templates/basic-skill/ (or templates/skill-with-scripts/ if you need scripts) to your chosen location.
    2. Rename the folder and set name and description.
    3. Run npm run qa in this guide’s folder if you’re developing inside it, or check the frontmatter by eye against the checklist.
  • Verify: Restart or reload the host. The skill appears when you type / in Cursor or Claude Code.

Step 6: Write the body (core workflow only)

Section titled “Step 6: Write the body (core workflow only)”
  • Goal: Concise instructions that close the gaps you found in Step 1.
  • Do:
    1. Start with a Quick start or Workflow section: numbered steps, in the imperative mood.
    2. Add only what the model doesn’t know, such as your categories, voice, forbidden words, and the output template.
    3. For multi-step work, include a copyable progress checklist (see the workflow pattern).
    4. Match the level of freedom to how fragile the task is (see degrees of freedom).
    5. Keep SKILL.md under 500 lines, and ideally much shorter.
  • Verify: Every paragraph answers “what would the agent get wrong without this?” If a paragraph doesn’t, delete it.
  • Goal: Keep level 2 small and push detail down to level 3.
  • Do:
    1. Move long reference material into references/*.md (style guide, glossary, API details).
    2. Put output templates in assets/.
    3. Link each file directly from SKILL.md with a sentence saying when to read it: For tone and banned phrases, read [references/voice.md](references/voice.md).
    4. Keep references one level deep. Don’t chain SKILL.md → a.md → b.md.
    5. Add a table of contents to any reference file over ~100 lines.
  • Verify: SKILL.md alone is enough for the common case; resources are only needed for the less common ones.

Step 8: Add scripts for deterministic work (optional)

Section titled “Step 8: Add scripts for deterministic work (optional)”
  • Goal: Reliability and token savings for fragile or mechanical steps.
  • Do:
    1. Write scripts that solve the problem and handle their own errors, and that print clear, actionable messages.
    2. Document each one in SKILL.md with its exact command and expected output, and say whether to run it or read it.
    3. Document dependencies (for example “requires Python 3.10+, pip install pypdf”).
    4. Use forward-slash paths (scripts/validate.py), never backslashes.
    5. Add a validation loop: run → fix → re-run until it passes.
  • Verify: Run each script by hand from the skill folder on sample input, and check it fails cleanly on bad input.
  • Goal: Evidence that it works. Don’t settle for a feeling.
  • Do: Follow Testing & evaluation:
    1. Should trigger: each Step 1 request (in a fresh chat) loads the skill.
    2. Should not trigger: 3+ near-miss requests don’t load it.
    3. Quality: compare outputs with your baseline. Is every gap from Step 1 closed?
    4. Explicit: /skill-name works.
    5. Cross-host: repeat on each host you support.
    6. Test with the models your users actually use. Smaller models may need more explicit steps.
  • Verify: The results are recorded in an evals/ file or table (see the template).
  • Goal: Improve based on what happens, not on guesses.
  • Do:
    1. Watch real sessions. Where does the agent skip steps, misread instructions, or read the wrong file?
    2. Fix the specific failure. Usually that means sharpening the description, moving something from a reference into the body (or the reverse), or adding an example.
    3. Re-run the eval set after every change.
  • Verify: The eval pass rate goes up and never regresses.
  • Goal: Others can use it with zero hand-holding.
  • Do: Follow Distribution. Commit it, add a changelog entry (in metadata or a CHANGELOG.md), and announce it with a one-line summary plus example prompts.
  • Verify: A colleague triggers it on their machine using only your announcement.

Next: Writing guide · Checklist