Agents Checklist
Copy this into your PR or ticket. Use Part A for subagents, Part B for SDK agents, and Part C for both.
Part A: Subagent (guide)
Section titled “Part A: Subagent (guide)”A1. Define
Section titled “A1. Define”- Job written in one sentence
- Input the parent must pass is defined
- Output shape and max length defined
- Stop condition defined
- 3–5 real tasks collected
A2. Justify
Section titled “A2. Justify”- Specific isolation benefit named (context, parallelism, permissions, model, independence)
- Not better served by a skill or rule
A3. Permissions & model
Section titled “A3. Permissions & model”- Least privilege: read-only unless editing is the job
- Claude Code:
tools:allowlist set (or deliberately omitted) - Cursor:
readonly:set appropriately - Model chosen (
inheritunless measured otherwise)
A4. Description
Section titled “A4. Description”- States when to use it; includes “use proactively” if automatic delegation is wanted
- No overlap with sibling agents
A5. System prompt
Section titled “A5. System prompt”- Role, first actions, prioritised checklist, constraints, output template, stop condition
- Self-sufficient: makes sense without the parent’s chat
A6. File
Section titled “A6. File”- Correct location (
.claude/agents/for both hosts,.cursor/agents/for Cursor only) - Listed in
/agentsor Customize → Agents
A7–A9. Test & tune
Section titled “A7–A9. Test & tune”- Explicit, implicit, and negative delegation tests pass
- Seeded-defect / golden tasks pass
- Output matches the template and length
- Constraints held (for example
git statusclean for a read-only agent) - Eval sheet saved
A10. Ship
Section titled “A10. Ship”- Committed or packaged
- README entry with example invocations
- Logged in
DECISIONS.md
Part B: SDK agent (guide)
Section titled “Part B: SDK agent (guide)”B1–B2. Define & choose
Section titled “B1–B2. Define & choose”- Trigger, inputs, output destination, and budget documented
- SDK choice (Claude / Cursor / headless CLI) and reason logged
B3. Setup
Section titled “B3. Setup”- Credentials in environment variables or CI secrets only; never committed
- “Hello” run succeeds
B4–B5. Build & restrict
Section titled “B4–B5. Build & restrict”- Minimal agent produces the output
- Tool allowlist set; permission mode is least permissive
- Runs in a sandbox or ephemeral runner with scoped tokens
- Out-of-scope request is refused or blocked
B6. Capabilities
Section titled “B6. Capabilities”- MCP servers / custom tools / setting sources configured deliberately (no accidental ambient config)
B7. Robustness
Section titled “B7. Robustness”- Structured output validated in code
-
maxTurns, timeout, and cost ceiling set - Startup failures and run failures handled separately, with distinct exit codes
- Retries only when the error is retryable, with backoff
- Resources disposed (
await using/with)
B8. Test
Section titled “B8. Test”- Golden task suite passes and runs in CI
- Each failure mode forced and handled
B9–B10. Deploy & operate
Section titled “B9–B10. Deploy & operate”- Trigger wired (CI / cron / webhook / cloud)
- Logs include prompt version, model, run ID, turns, cost, outcome
- Owner, monitoring, and failure alert in place
Part C: Both
Section titled “Part C: Both”- At least one hard guardrail on anything that changes state outside the repo
- Human checkpoint before irreversible actions
- Prompt-injection risk considered for any untrusted input
-
npm run qapasses - Versioned, with a changelog entry