Skip to content

MCP: Tool Design

The model decides which tool to call from its name, description, and schema alone. Tool design is prompt engineering, and it matters more than anything else in your server.

  1. Design for tasks, not endpoints. One schedule_meeting tool beats list_users + list_calendars + get_availability + create_event. Fewer calls means fewer chances for the model to go wrong.
  2. Keep the set small. Every tool’s schema occupies context in every request. Aim for under about 10 per server, and combine near-duplicates by adding a parameter.
  3. Return meaningful, concise output. Use human-readable names instead of opaque IDs (or both), and only the fields the model needs.
  4. Make errors instructive. “Date must be ISO 8601, e.g. 2026-09-29” gets fixed on the next try. “Invalid input” doesn’t.
  5. Be explicit about side effects with annotations and wording.
Rule ✅ ❌
verb_noun, snake_case search_orders, create_ticket orders, doTicketThing
Namespace when servers might collide jira_search_issues search
Consistent verbs across the server get_ / list_ / search_ / create_ / update_ / delete_ mixing fetch_, get_, retrieve_
Allowed characters letters, digits, _, - (keep under ~64 chars) spaces, dots

Write them for a smart colleague who has never seen your system:

Search customer orders by email, order number, or date range.
Returns at most 20 orders, newest first, with status and total.
Use get_order for full line items. Dates are ISO 8601 (YYYY-MM-DD).

Include:

  • What it does and what it returns
  • When to use it versus a sibling tool
  • Formats and limits (dates, units, pagination, maximum results)
  • Side effects (“sends an email to the customer”, “cannot be undone”)
  • Describe every parameter, including its format, units, and an example.
  • Use constraints: enum, minimum/maximum, minLength, pattern, and format: "date".
  • Keep the required parameters minimal and give the rest sensible defaults.
  • Prefer flat objects; deeply nested inputs are error-prone for models.
  • Use unambiguous parameter names: user_id, not user; amount_cents, not amount.
  • Text content that reads well: short lines and labelled fields.
  • Structured content (outputSchema + structuredContent) when a client or a follow-up tool needs to parse it. Include a text version too, for clients that only show text.
  • Truncate and paginate. Cap the result size and return a cursor or a “showing 20 of 312, refine with since” hint.
  • Resource links for large payloads: return a URI the client can fetch, instead of inlining megabytes.
  • Consider a response_format: "concise" | "detailed" parameter for tools whose output size varies a lot.
Annotation Meaning Set true for
readOnlyHint Doesn’t modify its environment search, get, list
destructiveHint May delete or overwrite (only meaningful when not read-only) delete, overwrite, bulk update
idempotentHint Repeating it with the same arguments has no additional effect upserts, “set status to X”
openWorldHint Interacts with external entities (web, third-party APIs) web fetch, sending email

Annotations are hints. Clients use them for approval UX, but they don’t enforce security. Real protection belongs in your server (see Security).

Situation Return
Bad input the model can fix isError: true + what’s wrong + a valid example
Not found isError: true + a suggestion (“No order 123. Use search_orders to find the number.”)
Permission denied isError: true + what access is needed (never leak secrets)
Upstream outage or timeout isError: true + “temporary, retry in N s”
Bug in your server Let it throw. The SDK returns an error; log the details to stderr
  • Total tool count per server is under about 10 (or the extras are split into optional servers)
  • Descriptions are complete but not padded (roughly 1–4 sentences)
  • Default result limits are small (5–20)
  • No raw HTML or huge JSON blobs in outputs
  • Large data is returned as resource links or paginated

A skill can teach the agent how to use your server well: which tool to use first, common workflows, and gotchas. This keeps tool descriptions short, while complex guidance loads only when it’s relevant. See Skills concepts (the “integration skill” archetype).