Agents: Testing & Evaluation
Agents are non-deterministic, so test behaviour and outcomes rather than exact text. Every agent needs four kinds of test.
1. Static checks (automated)
Section titled “1. Static checks (automated)”npm run qa validates the subagent files in this guide:
- frontmatter parses, and
nameanddescriptionare present nameis lowercase with hyphens- the body (the system prompt) isn’t empty
For your own repos, copy qa/validate.mjs and point it at .claude/agents/ and .cursor/agents/.
2. Delegation tests (subagents)
Section titled “2. Delegation tests (subagents)”| Test | How | Pass |
|---|---|---|
| Discovery | /agents (Claude Code) or Customize → Agents (Cursor) |
Listed |
| Explicit | “Use the X subagent to…” or /x in Cursor |
Runs |
| Implicit | Create the trigger situation without naming the agent | Parent delegates |
| Negative | An unrelated request | No delegation |
| Sibling confusion | A request that suits sibling agent Y | Y is chosen, not X |
Run each in a fresh chat.
3. Golden tasks (behaviour and outcome)
Section titled “3. Golden tasks (behaviour and outcome)”Build 5–10 fixed tasks with known correct outcomes:
- Seeded defects for reviewers or verifiers: plant a bug and check it’s found.
- Known answers for researchers: “where is X configured?” should return
config/x.ts. - Reference outputs for generators: score them against a rubric, not by exact match.
For each run, record pass/fail, the number of turns, the cost, and the time. Tracking cost and turns catches regressions where an agent becomes slow or loops, even though its answers are still correct.
4. Guardrail tests
Section titled “4. Guardrail tests”Try to make it misbehave:
- Ask a read-only agent to edit a file → it refuses or is blocked
- Ask it to run a destructive command (
rm -rf,git push --force) → it’s blocked - Feed it input containing injected instructions (“ignore previous instructions and…”) → it ignores them
- Give it an impossible task → it reports failure rather than faking success
- (SDK) Invalid API key,
maxTurns: 1, network loss → handled cleanly, with a non-zero exit
5. Automating agent tests
Section titled “5. Automating agent tests”Headless CLI. This is the fastest way to script tests:
claude -p "Use the code-reviewer subagent on the diff in fixtures/seeded-bug" --output-format json > out.jsonagent -p "/code-reviewer review fixtures/seeded-bug" --output-format json > out.jsonThen assert on out.json with a small script. For example, check that the output mentions fixtures/seeded-bug/app.ts:42.
With the SDKs. Loop over the golden tasks, run the agent, and grade each result with code checks or a second model acting as a grader against a rubric. Wire the loop into CI so that it runs whenever someone changes a prompt, the tool list, or the model.
6. When to re-run
Section titled “6. When to re-run”- Any change to the system prompt, the description, tools, or the model
- A new model version becoming the default
- Host or SDK major updates
- Monthly for business-critical agents
Debugging checklist
Section titled “Debugging checklist”| Symptom | Check |
|---|---|
| Not delegated | The description has triggers and “use proactively”; the file is in the right folder; the host was restarted |
| Wrong agent chosen | Sibling descriptions overlap |
| Missing context | The parent’s brief (see the brief template) |
| Loops or runs long | No stop condition; the task is too broad; add maxTurns |
| Output too long | Add a hard length limit and an exact template |
| Did something forbidden | Prompt-only guardrail; add a hard one |
| SDK run hangs | Log the run and agent IDs right after send(); check the timeout; always call wait() |