Skip to content

Agents Checklist

Copy this into your PR or ticket. Use Part A for subagents, Part B for SDK agents, and Part C for both.

  • Job written in one sentence
  • Input the parent must pass is defined
  • Output shape and max length defined
  • Stop condition defined
  • 3–5 real tasks collected
  • Specific isolation benefit named (context, parallelism, permissions, model, independence)
  • Not better served by a skill or rule
  • Least privilege: read-only unless editing is the job
  • Claude Code: tools: allowlist set (or deliberately omitted)
  • Cursor: readonly: set appropriately
  • Model chosen (inherit unless measured otherwise)
  • States when to use it; includes “use proactively” if automatic delegation is wanted
  • No overlap with sibling agents
  • Role, first actions, prioritised checklist, constraints, output template, stop condition
  • Self-sufficient: makes sense without the parent’s chat
  • Correct location (.claude/agents/ for both hosts, .cursor/agents/ for Cursor only)
  • Listed in /agents or Customize → Agents
  • Explicit, implicit, and negative delegation tests pass
  • Seeded-defect / golden tasks pass
  • Output matches the template and length
  • Constraints held (for example git status clean for a read-only agent)
  • Eval sheet saved
  • Committed or packaged
  • README entry with example invocations
  • Logged in DECISIONS.md
  • Trigger, inputs, output destination, and budget documented
  • SDK choice (Claude / Cursor / headless CLI) and reason logged
  • Credentials in environment variables or CI secrets only; never committed
  • “Hello” run succeeds
  • Minimal agent produces the output
  • Tool allowlist set; permission mode is least permissive
  • Runs in a sandbox or ephemeral runner with scoped tokens
  • Out-of-scope request is refused or blocked
  • MCP servers / custom tools / setting sources configured deliberately (no accidental ambient config)
  • Structured output validated in code
  • maxTurns, timeout, and cost ceiling set
  • Startup failures and run failures handled separately, with distinct exit codes
  • Retries only when the error is retryable, with backoff
  • Resources disposed (await using / with)
  • Golden task suite passes and runs in CI
  • Each failure mode forced and handled
  • Trigger wired (CI / cron / webhook / cloud)
  • Logs include prompt version, model, run ID, turns, cost, outcome
  • Owner, monitoring, and failure alert in place
  • At least one hard guardrail on anything that changes state outside the repo
  • Human checkpoint before irreversible actions
  • Prompt-injection risk considered for any untrusted input
  • npm run qa passes
  • Versioned, with a changelog entry