Agent SDK production deployment
Session models, isolation, observability, costs, structured output, tool search, cache TTLs and TypeScript/Python deployment examples.
Agent SDK embeds the Claude Code loop in services, jobs, CI and multitenant products. query() starts and supervises a claude CLI subprocess over stdio; it is not a stateless API request.
Each active session owns a process tree, workspace and local transcript. Plan filesystems, isolation, observability and costs as for stateful workers.
Read capabilities for permissions, hooks, MCP, SessionStore and costs; API reference for language APIs, sessions and migration; agent features for context sources, skills, subagents, tasks and checkpoints; and runtime patterns for tools, prompts, streaming, structured output, search, approvals and commands. This page focuses on topology, multitenancy and operations.
#Execution model
| Item | Production implication |
|---|---|
| query() | Starts a CLI subprocess rather than calling a stateless wrapper |
| Subprocess | Owns shell, tools, cwd and local session files |
| Concurrency | N sessions usually mean N subprocesses; budget CPU, memory, disk and API limits |
| cwd | Inherits the host directory by default; set explicitly for multiple sessions/tenants |
| Local state | Lost on restart, scaling or migration unless persisted separately |
There are three main categories of default local state:
| State | Default location | Production handling |
|---|---|---|
| Transcripts | ~/.claude/projects or projects/ under CLAUDE_CONFIG_DIR | Mirror with SessionStore for cross-host recovery |
| Memory | User ~/.claude/CLAUDE.md and project CLAUDE.md | Separate volumes, object storage or disabling policy; not replaced by SessionStore |
| Artifacts | Session cwd | Per-tenant/task directories, volumes or object-storage synchronization |
#Session models
| Model | Use case | Design |
|---|---|---|
| Ephemeral | One-shot fixes, analysis, conversion, CI | Container/sandbox per task; destroy afterwards and retain explicit exports only |
| Long-running | Slack bots, email agents, ongoing site builders | Persistent worker; route HTTP/WebSocket requests for a session to that worker |
| Hybrid + SessionStore | Intermittent research, project management and support | Release idle containers; restore transcripts by ID on return |
| Multi-agent container | Collaboration or simulation | Multiple subprocesses with separate cwd, config and permission boundaries |
SessionStore mirrors transcripts, not all local state. The subprocess writes locally before SDK batches reach the store. CLAUDE.md, auto memory, checkpoint blobs and workspace artifacts remain separate.
After bounded append retries fail, execution continues with a system/mirror_error event. Monitor it so missing transcript batches do not silently compromise restoration.
#Multitenant isolation
The SDK can load local user/project settings and memory. Without isolation, one tenant's CLAUDE.md, MCP, commands or auto memory may enter another tenant's context.
| Boundary | TypeScript | Python | Purpose |
|---|---|---|---|
| Filesystem settings | settingSources: [] | setting_sources=[] | Skip user/project/local settings |
| Auto memory | CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 | Same | Prevent project memory injection |
| Config directory | CLAUDE_CONFIG_DIR=/srv/claude-config/tenant | Same | Isolate config, transcripts and cached state |
| Workspace | cwd: tenantDir | cwd=tenant_dir | Separate files, commands and artifacts |
| Egress/proxy | Gateway/network configuration | Same | Independent egress, credentials, allowlists and auditing |
Older Python versions treated an empty setting_sources list as unset. Upgrade before relying on it for isolation.
#Observability and cost
A user task can span many model requests, tools, MCP calls and subagents. Export OpenTelemetry and record result token/cost fields.
CLAUDE_CODE_ENABLE_TELEMETRY=1
CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1
OTEL_TRACES_EXPORTER=otlp
OTEL_METRICS_EXPORTER=otlp
OTEL_LOGS_EXPORTER=otlp
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_EXPORTER_OTLP_ENDPOINT=http://collector.example.com:4318
OTEL_RESOURCE_ATTRIBUTES=service.name=agent-runtime,team=platformDo not log raw prompts, tool input/output or API bodies by default. Enable sensitive diagnostics briefly in an isolated environment only.
| Field | Meaning | Use |
|---|---|---|
| message.total_cost_usd | Client estimate | Per-session/tenant accounting and budget alerts |
| message.usage.input_tokens | Standard input tokens | Detect context growth |
| message.usage.output_tokens | Output tokens | Understand answer and tool-planning cost |
| message.usage.cache_creation_input_tokens | Cache writes | Usually priced above standard input |
| message.usage.cache_read_input_tokens | Cache reads | Measure savings at lower read rates |
Successful and failed result messages can both contain usage and costs. Record both.
#Structured output and tool scale
Use outputFormat/output_format for database, task-system or UI results. The result contains structured_output; the SDK validates the schema, retries mismatches and eventually returns an error on failure.
Custom tools and MCP are the main extension mechanisms:
| Capability | Usage |
|---|---|
| Custom tools | In-process MCP wraps functions, database access or internal APIs |
| Remote MCP | HTTP/SSE/stdio servers connect Slack, GitHub, databases and ticket systems |
| allowedTools / allowed_tools | Preapprove explicit tools or mcp__server__* patterns |
| Tool search | Defer definitions instead of putting every schema in each request |
Common ENABLE_TOOL_SEARCH values:
| Value | Behavior |
|---|---|
| Unset | Enabled by default; Vertex or non-first-party base URLs may fall back to upfront schemas |
| true | Force enable; unsupported tool_reference may fail |
| auto | Enable when definitions exceed 10% of context |
| auto:5 | Enable above 5% for earlier search |
| false | Load all definitions upfront |
For fewer than roughly ten tools, upfront loading is often simpler. Search can significantly reduce context and improve selection with dozens, hundreds or thousands.
#Prompt cache TTL
The SDK uses caching automatically. Usually no manual cache controls are needed, but understand five-minute versus one-hour costs.
| Rule | Explanation |
|---|---|
| API key / Bedrock / Vertex / Foundry | Writes typically default to a five-minute TTL |
| ENABLE_PROMPT_CACHING_1H=1 | Request one hour for repeated short sessions with stable context |
| Claude subscription | Usually one hour within plan allowance |
| One-hour writes | Higher write price; worthwhile with sufficient subsequent reads |
| Cache hit | Refreshes TTL; TTL measures inactivity, not total lifetime |
| Prefix changes | Model, effort, tools, MCP and compaction can reduce hits |
Inspect creation and read tokens separately. High creation every turn suggests changing prefixes; increasing reads indicate reuse.
#TypeScript configuration
import path from "node:path";
import { query, type SessionStore } from "@anthropic-ai/claude-agent-sdk";
declare const prompt: string;
declare const sessionId: string | undefined;
declare const sessionStore: SessionStore;
const tenantId = "tenant_123";
const tenantDir = path.join("/srv/agent-work", tenantId);
const configDir = path.join("/srv/claude-config", tenantId);
for await (const message of query({
prompt,
options: {
cwd: tenantDir,
resume: sessionId,
sessionStore,
maxTurns: 30,
maxBudgetUsd: 5,
permissionMode: "dontAsk",
settingSources: [],
allowedTools: ["Read", "Grep", "Glob", "mcp__enterprise-tools__*"],
mcpServers: {
"enterprise-tools": {
type: "http",
url: "https://tools.example.com/mcp"
}
},
outputFormat: {
type: "json_schema",
schema: {
type: "object",
properties: {
summary: { type: "string" },
actions: {
type: "array",
items: { type: "string" }
}
},
required: ["summary"]
}
},
env: {
...process.env,
CLAUDE_CONFIG_DIR: configDir,
CLAUDE_CODE_DISABLE_AUTO_MEMORY: "1",
CLAUDE_CODE_ENABLE_TELEMETRY: "1",
CLAUDE_CODE_ENHANCED_TELEMETRY_BETA: "1",
OTEL_TRACES_EXPORTER: "otlp",
OTEL_METRICS_EXPORTER: "otlp",
OTEL_LOGS_EXPORTER: "otlp",
OTEL_EXPORTER_OTLP_PROTOCOL: "http/protobuf",
OTEL_EXPORTER_OTLP_ENDPOINT: "http://collector.example.com:4318",
ENABLE_TOOL_SEARCH: "auto:5"
}
}
})) {
if (message.type === "system" && message.subtype === "mirror_error") {
console.error("SessionStore mirror failed", message);
}
if (message.type === "result") {
console.log({
subtype: message.subtype,
cost: message.total_cost_usd,
usage: message.usage,
structured: message.structured_output
});
}
}TypeScript env replaces rather than merges subprocess variables. Spread process.env to preserve PATH, keys and provider variables.
#Python configuration
import asyncio
from pathlib import Path
from claude_agent_sdk import ClaudeAgentOptions, query
session_store = ...
async def run_agent(prompt: str, session_id: str | None = None) -> None:
tenant_id = "tenant_123"
tenant_dir = Path("/srv/agent-work") / tenant_id
config_dir = Path("/srv/claude-config") / tenant_id
options = ClaudeAgentOptions(
cwd=tenant_dir,
resume=session_id,
session_store=session_store,
max_turns=30,
max_budget_usd=5,
permission_mode="dontAsk",
setting_sources=[],
allowed_tools=["Read", "Grep", "Glob", "mcp__enterprise-tools__*"],
mcp_servers={
"enterprise-tools": {
"type": "http",
"url": "https://tools.example.com/mcp",
}
},
output_format={
"type": "json_schema",
"schema": {
"type": "object",
"properties": {
"summary": {"type": "string"},
"actions": {
"type": "array",
"items": {"type": "string"},
},
},
"required": ["summary"],
},
},
env={
"CLAUDE_CONFIG_DIR": str(config_dir),
"CLAUDE_CODE_DISABLE_AUTO_MEMORY": "1",
"CLAUDE_CODE_ENABLE_TELEMETRY": "1",
"CLAUDE_CODE_ENHANCED_TELEMETRY_BETA": "1",
"OTEL_TRACES_EXPORTER": "otlp",
"OTEL_METRICS_EXPORTER": "otlp",
"OTEL_LOGS_EXPORTER": "otlp",
"OTEL_EXPORTER_OTLP_PROTOCOL": "http/protobuf",
"OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.example.com:4318",
"ENABLE_TOOL_SEARCH": "auto:5",
},
)
async for message in query(prompt=prompt, options=options):
if getattr(message, "type", None) == "system" and getattr(message, "subtype", None) == "mirror_error":
print("SessionStore mirror failed", message)
if getattr(message, "type", None) == "result":
print(
{
"subtype": getattr(message, "subtype", None),
"cost": getattr(message, "total_cost_usd", None),
"usage": getattr(message, "usage", None),
"structured": getattr(message, "structured_output", None),
}
)
asyncio.run(run_agent("Analyze this tenant workspace and return a structured summary"))Python env overlays inherited variables. Prefer container secrets or gateway injection for authentication and proxy settings.
#Launch checklist
| Check | Acceptance criterion |
|---|---|
| Process boundary | Explicit cwd, resources, turn limits and budget per session |
| Recovery | SessionStore configured where required; mirror_error monitored |
| Persistence | Separate retention policies for memory, artifacts and transcripts |
| Isolation | Empty setting sources, disabled auto memory, per-tenant config and egress implemented |
| Observability | OTEL traces/metrics/logs reach collector; sensitive logging off by default |
| Costs | Result costs, usage and cache reads/writes archived by tenant |
| Tools | Allowlisted tools/MCP; search evaluated for large inventories |
| Caching | Default five minutes or one-hour override chosen for session intervals |
#Official references
Support
Need help?
For setup, billing, or model issues, email us. Check the status page for uptime.
WeChat / QQ support is available at the bottom right.

