Skip to content

The Universal Build Lifecycle

Skills, agents, and MCP servers all follow the same seven phases. The section guides apply these phases to each one. This page explains why each phase exists and what “done” means.

1. Define → 2. Design → 3. Scaffold → 4. Build → 5. Test → 6. Ship → 7. Operate
▲ │
└──────────────────────── iterate on real usage ────────────────────┘

Phase 1: Define (the problem, not the solution)

Section titled “Phase 1: Define (the problem, not the solution)”
  • Goal: A one-paragraph problem statement and 3–5 concrete example requests you want handled.
  • Do:
    • Write down who uses it, what they ask, and what a good result looks like.
    • Collect real examples from chat history, tickets, or colleagues. Don’t invent them.
    • Run the examples without your extension first and note where the agent fails. That baseline tells you what the extension actually has to fix.
  • Verify: Someone else can read the statement and predict what the extension will do.

Phase 2: Design (the lightest mechanism that works)

Section titled “Phase 2: Design (the lightest mechanism that works)”
  • Goal: A chosen mechanism (see Choosing the right tool), a name, a description, and a scope.
  • Do:
    • Choose the mechanism, and write down in one sentence why the lighter options aren’t enough.
    • Draft the description first. If you can’t write a sharp one, the scope is wrong.
    • Decide the location: personal, project, team or plugin, or hosted.
    • List what’s in scope and, just as important, what’s explicitly out.
  • Verify: The description names both what it does and when to use it, and it doesn’t overlap with anything existing.
  • Goal: Files exist in the right place and the host detects them.
  • Do: Copy the matching template from this guide, rename it, and fill in the frontmatter.
  • Verify: The host lists it. That means the skill appears in / or settings, the subagent in /agents or the Customize panel, and the MCP server shows green with its tools listed.
  • Goal: It works for your example requests.
  • Do: Implement the smallest version that handles the examples. Add detail only when a test shows it’s missing.
  • Verify: Each Phase 1 example produces a good result.
  • Goal: Confidence that it triggers when it should, stays quiet when it shouldn’t, and gives correct results.
  • Do:
    • Trigger tests: prompts that should activate it, and near-miss prompts that shouldn’t.
    • Behaviour tests: each example request, checked against the expected result.
    • Edge cases: bad input, missing files, auth failures, huge outputs.
    • Cross-platform: try every host you claim to support (for example Claude Code and Cursor).
    • Automated: add it to npm run qa or your CI.
  • Verify: Every test passes, and they’re written down so they can be re-run.
  • Goal: Other people can install and use it without asking you anything.
  • Do: Version it, document install and usage in a README, commit it to the right location or publish it, and announce it.
  • Verify: A colleague installs it from your instructions alone and completes an example task.
  • Goal: It stays correct as models, platforms, and your codebase change.
  • Do:
    • Watch real usage: missed triggers, wrong triggers, failed tool calls.
    • Re-run the eval set when models or platforms update.
    • Log changes in DECISIONS.md or a changelog.
    • Retire it once the model handles the task well without help.
  • Verify: It has an owner, a changelog, and a last-reviewed date.
  • Problem statement and real example requests are written down
  • Baseline (without the extension) was measured
  • Description states what and when, in the third person
  • Lives in the correct location for its audience
  • Trigger, behaviour, and edge-case tests pass and are recorded
  • Tested on every host it claims to support
  • No secrets in files; credentials come from environment variables or a secret manager
  • README or usage notes exist; a colleague has installed it successfully
  • Has an owner, a version, and a changelog entry