Claude Code

Agent SDK runtime patterns

Runtime design for custom tools, system prompts, streaming, structured output, tool search, approvals and slash commands.

The SDK starts the Claude Code loop with tools, permissions, settings, CLAUDE.md, hooks and MCP. Deliberate runtime design prevents excess exposure, cache misses, stalled approvals and unreliable output handling.

Start with SDK basics. See API reference for APIs/settings/migration, agent features for loop/context/skills/subagents/checkpoints, and production deployment for isolation and observability.

#Runtime decisions

GoalCapabilityRisk
Call application functionsCustom tools and in-process MCPBroad schemas, thrown business errors, unrestricted tools
CLI-like behaviorclaude_code prompt presetReplacing safety and tool guidance
Realtime UI for long tasksStreaming outputIgnoring intermediate tools and failures
Write results into systemsStructured outputUnstable schemas and unhandled retry exhaustion
Hundreds of toolsTool searchProvider fallback without tool_reference
User approvalcanUseTool, hooks and deferStalled callbacks or prior automatic approval
SDK slash commandsDispatch supported commands such as /compactOnly system/init-listed commands are available

#Custom tools

Custom tools expose application capabilities through in-process MCP servers, suitable for internal APIs, databases, retrieval, tickets and business actions.

PartRequirement
NameUnique, clear, preferably verb plus object
DescriptionExplain when to call it; avoid vague run query descriptions
Input schemaZod in TypeScript; type dictionaries or JSON Schema in Python
HandlerReturn content, optionally structuredContent and isError
ServerWrap with createSdkMcpServer/create_sdk_mcp_server
ApprovalAdd to allowedTools for execution without prompts

Return isError: true for recoverable business failures rather than throwing every error and potentially interrupting the loop.

import { query, tool, createSdkMcpServer } from "@anthropic-ai/claude-agent-sdk";
import { z } from "zod";

const lookupInvoice = tool(
  "lookup_invoice",
  "Look up an invoice by ID and return billing status, amount, and due date",
  {
    invoiceId: z.string().describe("Invoice ID, for example inv_123")
  },
  async ({ invoiceId }) => {
    const invoice = await getInvoice(invoiceId);
    if (!invoice) {
      return {
        isError: true,
        content: [{ type: "text", text: "Invoice not found" }]
      };
    }
    return {
      content: [{ type: "text", text: `Status: ${invoice.status}` }],
      structuredContent: invoice
    };
  }
);

const billingServer = createSdkMcpServer({
  name: "billing",
  version: "1.0.0",
  tools: [lookupInvoice]
});

for await (const message of query({
  prompt: "Check invoice inv_123 and summarize the next action",
  options: {
    mcpServers: { billing: billingServer },
    allowedTools: ["mcp__billing__lookup_invoice"]
  }
})) {
  console.log(message);
}

#Choose a system prompt

ChoiceUse caseRetains Claude Code behavior
UnsetMinimal tool loopNo, only basic tool support
claude_code presetCLI/IDE coding agentYes
Preset plus appendCoding agent with product rulesYes; lowest risk
Custom stringNoncoding role or permission modelNo; provide complete tool and safety instructions yourself

For coding agents, prefer the preset plus append:

{
  systemPrompt: {
    type: "preset",
    preset: "claude_code",
    append: "When changing billing code, always mention migration and rollback risks."
  },
  settingSources: ["project"]
}

CLAUDE.md is injected project context, not the system prompt, and works with any prompt. Use settingSources: [] to avoid loading host rules in multitenant services.

#Streaming output

Streaming suits terminals, WebSockets, dashboards and log UIs. Do not wait only for the final result.

MessageUI handling
system/initRecord session ID, commands, model and permission mode
Assistant textAppend to the visible stream
Tool useShow file reading, commands or MCP work
Tool resultSummarize; collapse sensitive details by default
Successful resultFinish and record cost/usage
Error resultShow recoverable failure and preserve transcript

Single input suits simple tasks. Continuous chat requires explicit cancellation, timeout and navigation-away handling.

#Structured output

Structured outputs validate final results with JSON Schema, Zod or Pydantic for databases and UIs.

Design areaRecommendation
SchemaSmall, stable fields; avoid deeply arbitrary objects
FallbackReturn error state rather than writing partial results
ProseKeep explanations in result text or a separate field
RetryAllow SDK retries within turn/budget limits
AuditSave schema versions for migration
outputFormat: {
  type: "json_schema",
  schema: {
    type: "object",
    properties: {
      risk: { enum: ["low", "medium", "high"] },
      summary: { type: "string" },
      actions: {
        type: "array",
        items: { type: "string" }
      }
    },
    required: ["risk", "summary", "actions"]
  }
}

Tool definitions consume context. Search large catalogs first and load the most relevant tools.

ValueBehavior
ENABLE_TOOL_SEARCH=trueForce enable
ENABLE_TOOL_SEARCH=autoEnable above about 10% of context
ENABLE_TOOL_SEARCH=auto:5Enable above about 5%
ENABLE_TOOL_SEARCH=falseDisable and load all definitions

Names and descriptions directly affect search quality:

WeakStrong
querysearch_slack_messages
Get dataSearch Slack messages by keyword, channel, user, or date range
runcreate_jira_ticket

On Vertex, Bedrock, Foundry or custom gateways, verify tool-reference support. Upfront fallback changes context usage and caching.

#User approvals and questions

canUseTool handles two kinds of pause:

TriggerMeaning
Tool approvalAn action was not preapproved
AskUserQuestionThe agent needs a choice or clarification

Earlier allow rules, modes or allowedTools can skip canUseTool. Put mandatory checks in PreToolUse.

SituationHandling
User onlineShow approval UI and return allow/deny
User possibly offlineDefer, persist and resume later
High-risk actionBlock in a hook or require human review
External notificationPermissionRequest hook can notify Slack/email/push

#Slash commands in SDK

Send noninteractive slash commands as prompts. Discover supported commands from system/init slash_commands.

CommandSDK useCache effect
/compactCompress long historyRewrites prefix; next turn may miss
/contextDiagnose context usageSmall impact; mainly another message
/usageDisplay usage and cachingSmall impact; useful for dashboards
/clearUsually unsuitable for continuing sessionsDiscards context and starts a new prefix
Bundled skillsDepends on available commandsSkill prompt appends to conversation

/compact requires existing history; invoking it immediately in a new one-turn task is not useful.

#Prompt-cache TTL strategy

Caching is automatic. Hits depend on stable model, effort, prompts, tools, MCP, settings, CLAUDE.md and history prefixes.

ScenarioTTL recommendation
One-shot tasks, CI, temporary fixesFive minutes is usually sufficient
Frequent workspace conversationsOptional one-hour TTL may pay off
Many upfront toolsPrefer tool search before extending TTL
Frequently changing prompts/toolsOne hour still cannot fix unstable prefixes
Claude subscriptionsUsually one hour within the plan

Default behavior uses five minutes. Request one hour when reusing the same prefix after longer intervals:

export ENABLE_PROMPT_CACHING_1H=1

Observe separately:

FieldMeaning
cache_creation_input_tokensWrites; high values indicate large or unstable prefixes
cache_read_input_tokensReads; high values indicate effective reuse
input_tokensStandard uncached input
total_cost_usdEstimate; record successes and failures

#Official references

Support

Need help?

For setup, billing, or model issues, email us. Check the status page for uptime.

WeChat / QQ support is available at the bottom right.