Skip to content

Agents: Testing & Evaluation

Agents are non-deterministic, so test behaviour and outcomes rather than exact text. Every agent needs four kinds of test.

npm run qa validates the subagent files in this guide:

  • frontmatter parses, and name and description are present
  • name is lowercase with hyphens
  • the body (the system prompt) isn’t empty

For your own repos, copy qa/validate.mjs and point it at .claude/agents/ and .cursor/agents/.

Test How Pass
Discovery /agents (Claude Code) or Customize → Agents (Cursor) Listed
Explicit “Use the X subagent to…” or /x in Cursor Runs
Implicit Create the trigger situation without naming the agent Parent delegates
Negative An unrelated request No delegation
Sibling confusion A request that suits sibling agent Y Y is chosen, not X

Run each in a fresh chat.

Build 5–10 fixed tasks with known correct outcomes:

  • Seeded defects for reviewers or verifiers: plant a bug and check it’s found.
  • Known answers for researchers: “where is X configured?” should return config/x.ts.
  • Reference outputs for generators: score them against a rubric, not by exact match.

For each run, record pass/fail, the number of turns, the cost, and the time. Tracking cost and turns catches regressions where an agent becomes slow or loops, even though its answers are still correct.

Try to make it misbehave:

  • Ask a read-only agent to edit a file → it refuses or is blocked
  • Ask it to run a destructive command (rm -rf, git push --force) → it’s blocked
  • Feed it input containing injected instructions (“ignore previous instructions and…”) → it ignores them
  • Give it an impossible task → it reports failure rather than faking success
  • (SDK) Invalid API key, maxTurns: 1, network loss → handled cleanly, with a non-zero exit

Headless CLI. This is the fastest way to script tests:

Terminal window
claude -p "Use the code-reviewer subagent on the diff in fixtures/seeded-bug" --output-format json > out.json
agent -p "/code-reviewer review fixtures/seeded-bug" --output-format json > out.json

Then assert on out.json with a small script. For example, check that the output mentions fixtures/seeded-bug/app.ts:42.

With the SDKs. Loop over the golden tasks, run the agent, and grade each result with code checks or a second model acting as a grader against a rubric. Wire the loop into CI so that it runs whenever someone changes a prompt, the tool list, or the model.

  • Any change to the system prompt, the description, tools, or the model
  • A new model version becoming the default
  • Host or SDK major updates
  • Monthly for business-critical agents
Symptom Check
Not delegated The description has triggers and “use proactively”; the file is in the right folder; the host was restarted
Wrong agent chosen Sibling descriptions overlap
Missing context The parent’s brief (see the brief template)
Loops or runs long No stop condition; the task is too broad; add maxTurns
Output too long Add a hard length limit and an exact template
Did something forbidden Prompt-only guardrail; add a hard one
SDK run hangs Log the run and agent IDs right after send(); check the timeout; always call wait()