Skip to content

Step-by-Step: Build an SDK Agent (Claude Agent SDK & Cursor SDK)

Use this guide when an agent must run from code: a CI job, a bot, a scheduled task, or inside a product. Each step has a Goal, what to Do, and how to Verify it.

Running example: a pr-summary agent that runs in CI, reads the diff, and writes a summary to summary.md.

SDK APIs change faster than the file formats. The snippets here show the shape of each SDK. Always check them against the current docs in the official sources before you rely on them.


Step 1: Define the job, trigger, and output

Section titled “Step 1: Define the job, trigger, and output”
  • Goal: Know what starts the agent, what it does, and where its result goes.
  • Do: Write down:
    • Trigger: a PR is opened (CI), a cron schedule, a webhook, or a user action in your app
    • Inputs: the repo checkout, the PR number, environment variables
    • Output: a file, a PR comment, a JSON result, or an exit code
    • Budget: maximum turns, maximum cost, timeout
  • Verify: You could implement it as a shell script if the “thinking” part were done by a human.
  • Goal: The right harness.

  • Do:

    Choose If
    Claude Agent SDK You want to run on your own infrastructure with Anthropic models (or via Bedrock/Vertex), and need fine-grained tool, permission, and hook control
    Cursor SDK You want the Cursor agent (any model Cursor offers), your existing .cursor/ config, and optionally Cursor Cloud Agents to run it remotely
    Headless CLI (claude -p / agent -p) A one-off in a shell script or CI step is enough. It’s the simplest option
  • Verify: The choice and the reason are written in DECISIONS.md.

Step 3: Set up the project and credentials

Section titled “Step 3: Set up the project and credentials”
  • Goal: A runnable skeleton that authenticates.

  • Do:

    Terminal window
    # Claude Agent SDK (TypeScript)
    npm install @anthropic-ai/claude-agent-sdk
    # Python: pip install claude-agent-sdk
    # Auth: set ANTHROPIC_API_KEY
    # Cursor SDK (TypeScript, Node >= 22.13)
    npm install @cursor/sdk
    # Python: pip install cursor-sdk
    # Auth: set CURSOR_API_KEY

    Keep keys in environment variables or CI secrets. Never commit them.

  • Verify: A “hello” prompt runs and returns text.

  • Goal: The smallest agent that does the job.

  • Do (Claude Agent SDK, TypeScript): see templates/claude-sdk-agent.ts:

    import { query } from "@anthropic-ai/claude-agent-sdk";
    for await (const message of query({
    prompt: "Summarise the changes in `git diff origin/main...HEAD` into summary.md",
    options: {
    systemPrompt: "You are a release engineer. Be concise and factual.",
    allowedTools: ["Read", "Grep", "Glob", "Bash", "Write"],
    permissionMode: "acceptEdits",
    maxTurns: 20,
    },
    })) {
    if (message.type === "result") console.log(message.subtype, message.total_cost_usd);
    }
  • Do (Cursor SDK, TypeScript): see templates/cursor-sdk-agent.ts. The flow is: create an agent, send a prompt, then stream or wait for the run.

  • Verify: It runs locally against a real repo and produces the output.

  • Goal: Least privilege for unattended runs.
  • Do:
    1. Allowlist only the tools it needs (allowedTools, disallowedTools).
    2. Pick the least permissive permission mode that still works. Avoid bypassing permissions except in a disposable sandbox.
    3. Run in a container or ephemeral CI runner with a scoped token (for example, a PR-comment-only GitHub token).
    4. Add hooks to block dangerous commands if it has shell access.
  • Verify: Try to make it do something outside its job (for example “delete the README”). It refuses or is blocked.

Step 6: Add tools and knowledge (optional)

Section titled “Step 6: Add tools and knowledge (optional)”
  • Goal: Give it the capabilities and procedures it needs.
  • Do:
    • MCP servers: pass them in the options (mcpServers), or load project config (Claude: settingSources: ["project"]; Cursor local: local.settingSources).
    • Custom in-process tools (Claude SDK): define them with tool() + createSdkMcpServer(). This is useful for calling your app’s own functions.
    • Skills and subagents: load them from the project by enabling setting sources, or define subagents inline with the agents option (Claude SDK).
  • Verify: The agent lists and uses the new tools on a test task.
  • Goal: Predictable behaviour in automation.
  • Do:
    1. Ask for structured output (JSON matching a schema, or a fixed Markdown template) and validate it in code.
    2. Set maxTurns, a timeout, and a cost ceiling. Log the cost per run.
    3. Handle errors: an auth failure, a rate limit (retry with backoff), a max-turns stop, or malformed output (retry once, then fail loudly).
    4. Return a non-zero exit code when it fails, so CI notices.
  • Verify: Force each failure mode (a bad key, maxTurns: 1, a malformed prompt) and confirm it’s handled cleanly.
  • Goal: Regression protection.
  • Do: Build 5–10 fixed tasks with known good outcomes (for example, sample diffs with expected summary points). Run them on every change to the prompt, tools, or model, and assert on the results. See Testing & evaluation.
  • Verify: The golden suite passes and is wired into CI.
  • Goal: It runs on its trigger without you.
  • Do:
    • CI: add a job that installs the SDK, sets the secret, and runs the script. Use --output-format json for CLI variants.
    • Scheduled: a cron job or scheduled workflow.
    • Cursor Cloud Agents: run it on Cursor’s infrastructure with the SDK’s cloud runtime or the REST API.
    • Add logging: the prompt version, model, turns, cost, and outcome.
  • Verify: Trigger it for real (open a test PR). The output appears, and the logs are complete.
  • Goal: It stays healthy.
  • Do: Monitor the failure rate and cost, re-run the golden tasks when models update, version the system prompt, and log changes in DECISIONS.md.
  • Verify: There’s an owner, a dashboard or log query, and an alert on repeated failures.

Next: Design patterns · Checklist