Agent SDK runtime patterns
Runtime design for custom tools, system prompts, streaming, structured output, tool search, approvals and slash commands.
The SDK starts the Claude Code loop with tools, permissions, settings, CLAUDE.md, hooks and MCP. Deliberate runtime design prevents excess exposure, cache misses, stalled approvals and unreliable output handling.
Start with SDK basics. See API reference for APIs/settings/migration, agent features for loop/context/skills/subagents/checkpoints, and production deployment for isolation and observability.
#Runtime decisions
| Goal | Capability | Risk |
|---|---|---|
| Call application functions | Custom tools and in-process MCP | Broad schemas, thrown business errors, unrestricted tools |
| CLI-like behavior | claude_code prompt preset | Replacing safety and tool guidance |
| Realtime UI for long tasks | Streaming output | Ignoring intermediate tools and failures |
| Write results into systems | Structured output | Unstable schemas and unhandled retry exhaustion |
| Hundreds of tools | Tool search | Provider fallback without tool_reference |
| User approval | canUseTool, hooks and defer | Stalled callbacks or prior automatic approval |
| SDK slash commands | Dispatch supported commands such as /compact | Only system/init-listed commands are available |
#Custom tools
Custom tools expose application capabilities through in-process MCP servers, suitable for internal APIs, databases, retrieval, tickets and business actions.
| Part | Requirement |
|---|---|
| Name | Unique, clear, preferably verb plus object |
| Description | Explain when to call it; avoid vague run query descriptions |
| Input schema | Zod in TypeScript; type dictionaries or JSON Schema in Python |
| Handler | Return content, optionally structuredContent and isError |
| Server | Wrap with createSdkMcpServer/create_sdk_mcp_server |
| Approval | Add to allowedTools for execution without prompts |
Return isError: true for recoverable business failures rather than throwing every error and potentially interrupting the loop.
import { query, tool, createSdkMcpServer } from "@anthropic-ai/claude-agent-sdk";
import { z } from "zod";
const lookupInvoice = tool(
"lookup_invoice",
"Look up an invoice by ID and return billing status, amount, and due date",
{
invoiceId: z.string().describe("Invoice ID, for example inv_123")
},
async ({ invoiceId }) => {
const invoice = await getInvoice(invoiceId);
if (!invoice) {
return {
isError: true,
content: [{ type: "text", text: "Invoice not found" }]
};
}
return {
content: [{ type: "text", text: `Status: ${invoice.status}` }],
structuredContent: invoice
};
}
);
const billingServer = createSdkMcpServer({
name: "billing",
version: "1.0.0",
tools: [lookupInvoice]
});
for await (const message of query({
prompt: "Check invoice inv_123 and summarize the next action",
options: {
mcpServers: { billing: billingServer },
allowedTools: ["mcp__billing__lookup_invoice"]
}
})) {
console.log(message);
}#Choose a system prompt
| Choice | Use case | Retains Claude Code behavior |
|---|---|---|
| Unset | Minimal tool loop | No, only basic tool support |
| claude_code preset | CLI/IDE coding agent | Yes |
| Preset plus append | Coding agent with product rules | Yes; lowest risk |
| Custom string | Noncoding role or permission model | No; provide complete tool and safety instructions yourself |
For coding agents, prefer the preset plus append:
{
systemPrompt: {
type: "preset",
preset: "claude_code",
append: "When changing billing code, always mention migration and rollback risks."
},
settingSources: ["project"]
}CLAUDE.md is injected project context, not the system prompt, and works with any prompt. Use settingSources: [] to avoid loading host rules in multitenant services.
#Streaming output
Streaming suits terminals, WebSockets, dashboards and log UIs. Do not wait only for the final result.
| Message | UI handling |
|---|---|
| system/init | Record session ID, commands, model and permission mode |
| Assistant text | Append to the visible stream |
| Tool use | Show file reading, commands or MCP work |
| Tool result | Summarize; collapse sensitive details by default |
| Successful result | Finish and record cost/usage |
| Error result | Show recoverable failure and preserve transcript |
Single input suits simple tasks. Continuous chat requires explicit cancellation, timeout and navigation-away handling.
#Structured output
Structured outputs validate final results with JSON Schema, Zod or Pydantic for databases and UIs.
| Design area | Recommendation |
|---|---|
| Schema | Small, stable fields; avoid deeply arbitrary objects |
| Fallback | Return error state rather than writing partial results |
| Prose | Keep explanations in result text or a separate field |
| Retry | Allow SDK retries within turn/budget limits |
| Audit | Save schema versions for migration |
outputFormat: {
type: "json_schema",
schema: {
type: "object",
properties: {
risk: { enum: ["low", "medium", "high"] },
summary: { type: "string" },
actions: {
type: "array",
items: { type: "string" }
}
},
required: ["risk", "summary", "actions"]
}
}#Tool search
Tool definitions consume context. Search large catalogs first and load the most relevant tools.
| Value | Behavior |
|---|---|
| ENABLE_TOOL_SEARCH=true | Force enable |
| ENABLE_TOOL_SEARCH=auto | Enable above about 10% of context |
| ENABLE_TOOL_SEARCH=auto:5 | Enable above about 5% |
| ENABLE_TOOL_SEARCH=false | Disable and load all definitions |
Names and descriptions directly affect search quality:
| Weak | Strong |
|---|---|
query | search_slack_messages |
Get data | Search Slack messages by keyword, channel, user, or date range |
run | create_jira_ticket |
On Vertex, Bedrock, Foundry or custom gateways, verify tool-reference support. Upfront fallback changes context usage and caching.
#User approvals and questions
canUseTool handles two kinds of pause:
| Trigger | Meaning |
|---|---|
| Tool approval | An action was not preapproved |
| AskUserQuestion | The agent needs a choice or clarification |
Earlier allow rules, modes or allowedTools can skip canUseTool. Put mandatory checks in PreToolUse.
| Situation | Handling |
|---|---|
| User online | Show approval UI and return allow/deny |
| User possibly offline | Defer, persist and resume later |
| High-risk action | Block in a hook or require human review |
| External notification | PermissionRequest hook can notify Slack/email/push |
#Slash commands in SDK
Send noninteractive slash commands as prompts. Discover supported commands from system/init slash_commands.
| Command | SDK use | Cache effect |
|---|---|---|
| /compact | Compress long history | Rewrites prefix; next turn may miss |
| /context | Diagnose context usage | Small impact; mainly another message |
| /usage | Display usage and caching | Small impact; useful for dashboards |
| /clear | Usually unsuitable for continuing sessions | Discards context and starts a new prefix |
| Bundled skills | Depends on available commands | Skill prompt appends to conversation |
/compact requires existing history; invoking it immediately in a new one-turn task is not useful.
#Prompt-cache TTL strategy
Caching is automatic. Hits depend on stable model, effort, prompts, tools, MCP, settings, CLAUDE.md and history prefixes.
| Scenario | TTL recommendation |
|---|---|
| One-shot tasks, CI, temporary fixes | Five minutes is usually sufficient |
| Frequent workspace conversations | Optional one-hour TTL may pay off |
| Many upfront tools | Prefer tool search before extending TTL |
| Frequently changing prompts/tools | One hour still cannot fix unstable prefixes |
| Claude subscriptions | Usually one hour within the plan |
Default behavior uses five minutes. Request one hour when reusing the same prefix after longer intervals:
export ENABLE_PROMPT_CACHING_1H=1Observe separately:
| Field | Meaning |
|---|---|
| cache_creation_input_tokens | Writes; high values indicate large or unstable prefixes |
| cache_read_input_tokens | Reads; high values indicate effective reuse |
| input_tokens | Standard uncached input |
| total_cost_usd | Estimate; record successes and failures |
#Official references
- Custom tools
- Modifying system prompts
- Streaming output
- Streaming Input
- Plugins in the SDK
- Structured outputs
- Tool search
- User input
- Slash commands in the SDK
#Related pages
Support
Need help?
For setup, billing, or model issues, email us. Check the status page for uptime.
WeChat / QQ support is available at the bottom right.

