Claude Code

Agent SDK production deployment

Session models, isolation, observability, costs, structured output, tool search, cache TTLs and TypeScript/Python deployment examples.

Agent SDK embeds the Claude Code loop in services, jobs, CI and multitenant products. query() starts and supervises a claude CLI subprocess over stdio; it is not a stateless API request.

Each active session owns a process tree, workspace and local transcript. Plan filesystems, isolation, observability and costs as for stateful workers.

Read capabilities for permissions, hooks, MCP, SessionStore and costs; API reference for language APIs, sessions and migration; agent features for context sources, skills, subagents, tasks and checkpoints; and runtime patterns for tools, prompts, streaming, structured output, search, approvals and commands. This page focuses on topology, multitenancy and operations.

#Execution model

ItemProduction implication
query()Starts a CLI subprocess rather than calling a stateless wrapper
SubprocessOwns shell, tools, cwd and local session files
ConcurrencyN sessions usually mean N subprocesses; budget CPU, memory, disk and API limits
cwdInherits the host directory by default; set explicitly for multiple sessions/tenants
Local stateLost on restart, scaling or migration unless persisted separately

There are three main categories of default local state:

StateDefault locationProduction handling
Transcripts~/.claude/projects or projects/ under CLAUDE_CONFIG_DIRMirror with SessionStore for cross-host recovery
MemoryUser ~/.claude/CLAUDE.md and project CLAUDE.mdSeparate volumes, object storage or disabling policy; not replaced by SessionStore
ArtifactsSession cwdPer-tenant/task directories, volumes or object-storage synchronization

#Session models

ModelUse caseDesign
EphemeralOne-shot fixes, analysis, conversion, CIContainer/sandbox per task; destroy afterwards and retain explicit exports only
Long-runningSlack bots, email agents, ongoing site buildersPersistent worker; route HTTP/WebSocket requests for a session to that worker
Hybrid + SessionStoreIntermittent research, project management and supportRelease idle containers; restore transcripts by ID on return
Multi-agent containerCollaboration or simulationMultiple subprocesses with separate cwd, config and permission boundaries

SessionStore mirrors transcripts, not all local state. The subprocess writes locally before SDK batches reach the store. CLAUDE.md, auto memory, checkpoint blobs and workspace artifacts remain separate.

After bounded append retries fail, execution continues with a system/mirror_error event. Monitor it so missing transcript batches do not silently compromise restoration.

#Multitenant isolation

The SDK can load local user/project settings and memory. Without isolation, one tenant's CLAUDE.md, MCP, commands or auto memory may enter another tenant's context.

BoundaryTypeScriptPythonPurpose
Filesystem settingssettingSources: []setting_sources=[]Skip user/project/local settings
Auto memoryCLAUDE_CODE_DISABLE_AUTO_MEMORY=1SamePrevent project memory injection
Config directoryCLAUDE_CONFIG_DIR=/srv/claude-config/tenantSameIsolate config, transcripts and cached state
Workspacecwd: tenantDircwd=tenant_dirSeparate files, commands and artifacts
Egress/proxyGateway/network configurationSameIndependent egress, credentials, allowlists and auditing

Older Python versions treated an empty setting_sources list as unset. Upgrade before relying on it for isolation.

#Observability and cost

A user task can span many model requests, tools, MCP calls and subagents. Export OpenTelemetry and record result token/cost fields.

CLAUDE_CODE_ENABLE_TELEMETRY=1
CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1
OTEL_TRACES_EXPORTER=otlp
OTEL_METRICS_EXPORTER=otlp
OTEL_LOGS_EXPORTER=otlp
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_EXPORTER_OTLP_ENDPOINT=http://collector.example.com:4318
OTEL_RESOURCE_ATTRIBUTES=service.name=agent-runtime,team=platform

Do not log raw prompts, tool input/output or API bodies by default. Enable sensitive diagnostics briefly in an isolated environment only.

FieldMeaningUse
message.total_cost_usdClient estimatePer-session/tenant accounting and budget alerts
message.usage.input_tokensStandard input tokensDetect context growth
message.usage.output_tokensOutput tokensUnderstand answer and tool-planning cost
message.usage.cache_creation_input_tokensCache writesUsually priced above standard input
message.usage.cache_read_input_tokensCache readsMeasure savings at lower read rates

Successful and failed result messages can both contain usage and costs. Record both.

#Structured output and tool scale

Use outputFormat/output_format for database, task-system or UI results. The result contains structured_output; the SDK validates the schema, retries mismatches and eventually returns an error on failure.

Custom tools and MCP are the main extension mechanisms:

CapabilityUsage
Custom toolsIn-process MCP wraps functions, database access or internal APIs
Remote MCPHTTP/SSE/stdio servers connect Slack, GitHub, databases and ticket systems
allowedTools / allowed_toolsPreapprove explicit tools or mcp__server__* patterns
Tool searchDefer definitions instead of putting every schema in each request

Common ENABLE_TOOL_SEARCH values:

ValueBehavior
UnsetEnabled by default; Vertex or non-first-party base URLs may fall back to upfront schemas
trueForce enable; unsupported tool_reference may fail
autoEnable when definitions exceed 10% of context
auto:5Enable above 5% for earlier search
falseLoad all definitions upfront

For fewer than roughly ten tools, upfront loading is often simpler. Search can significantly reduce context and improve selection with dozens, hundreds or thousands.

#Prompt cache TTL

The SDK uses caching automatically. Usually no manual cache controls are needed, but understand five-minute versus one-hour costs.

RuleExplanation
API key / Bedrock / Vertex / FoundryWrites typically default to a five-minute TTL
ENABLE_PROMPT_CACHING_1H=1Request one hour for repeated short sessions with stable context
Claude subscriptionUsually one hour within plan allowance
One-hour writesHigher write price; worthwhile with sufficient subsequent reads
Cache hitRefreshes TTL; TTL measures inactivity, not total lifetime
Prefix changesModel, effort, tools, MCP and compaction can reduce hits

Inspect creation and read tokens separately. High creation every turn suggests changing prefixes; increasing reads indicate reuse.

#TypeScript configuration

import path from "node:path";
import { query, type SessionStore } from "@anthropic-ai/claude-agent-sdk";

declare const prompt: string;
declare const sessionId: string | undefined;
declare const sessionStore: SessionStore;

const tenantId = "tenant_123";
const tenantDir = path.join("/srv/agent-work", tenantId);
const configDir = path.join("/srv/claude-config", tenantId);

for await (const message of query({
  prompt,
  options: {
    cwd: tenantDir,
    resume: sessionId,
    sessionStore,
    maxTurns: 30,
    maxBudgetUsd: 5,
    permissionMode: "dontAsk",
    settingSources: [],
    allowedTools: ["Read", "Grep", "Glob", "mcp__enterprise-tools__*"],
    mcpServers: {
      "enterprise-tools": {
        type: "http",
        url: "https://tools.example.com/mcp"
      }
    },
    outputFormat: {
      type: "json_schema",
      schema: {
        type: "object",
        properties: {
          summary: { type: "string" },
          actions: {
            type: "array",
            items: { type: "string" }
          }
        },
        required: ["summary"]
      }
    },
    env: {
      ...process.env,
      CLAUDE_CONFIG_DIR: configDir,
      CLAUDE_CODE_DISABLE_AUTO_MEMORY: "1",
      CLAUDE_CODE_ENABLE_TELEMETRY: "1",
      CLAUDE_CODE_ENHANCED_TELEMETRY_BETA: "1",
      OTEL_TRACES_EXPORTER: "otlp",
      OTEL_METRICS_EXPORTER: "otlp",
      OTEL_LOGS_EXPORTER: "otlp",
      OTEL_EXPORTER_OTLP_PROTOCOL: "http/protobuf",
      OTEL_EXPORTER_OTLP_ENDPOINT: "http://collector.example.com:4318",
      ENABLE_TOOL_SEARCH: "auto:5"
    }
  }
})) {
  if (message.type === "system" && message.subtype === "mirror_error") {
    console.error("SessionStore mirror failed", message);
  }

  if (message.type === "result") {
    console.log({
      subtype: message.subtype,
      cost: message.total_cost_usd,
      usage: message.usage,
      structured: message.structured_output
    });
  }
}

TypeScript env replaces rather than merges subprocess variables. Spread process.env to preserve PATH, keys and provider variables.

#Python configuration

import asyncio
from pathlib import Path

from claude_agent_sdk import ClaudeAgentOptions, query

session_store = ...


async def run_agent(prompt: str, session_id: str | None = None) -> None:
    tenant_id = "tenant_123"
    tenant_dir = Path("/srv/agent-work") / tenant_id
    config_dir = Path("/srv/claude-config") / tenant_id

    options = ClaudeAgentOptions(
        cwd=tenant_dir,
        resume=session_id,
        session_store=session_store,
        max_turns=30,
        max_budget_usd=5,
        permission_mode="dontAsk",
        setting_sources=[],
        allowed_tools=["Read", "Grep", "Glob", "mcp__enterprise-tools__*"],
        mcp_servers={
            "enterprise-tools": {
                "type": "http",
                "url": "https://tools.example.com/mcp",
            }
        },
        output_format={
            "type": "json_schema",
            "schema": {
                "type": "object",
                "properties": {
                    "summary": {"type": "string"},
                    "actions": {
                        "type": "array",
                        "items": {"type": "string"},
                    },
                },
                "required": ["summary"],
            },
        },
        env={
            "CLAUDE_CONFIG_DIR": str(config_dir),
            "CLAUDE_CODE_DISABLE_AUTO_MEMORY": "1",
            "CLAUDE_CODE_ENABLE_TELEMETRY": "1",
            "CLAUDE_CODE_ENHANCED_TELEMETRY_BETA": "1",
            "OTEL_TRACES_EXPORTER": "otlp",
            "OTEL_METRICS_EXPORTER": "otlp",
            "OTEL_LOGS_EXPORTER": "otlp",
            "OTEL_EXPORTER_OTLP_PROTOCOL": "http/protobuf",
            "OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.example.com:4318",
            "ENABLE_TOOL_SEARCH": "auto:5",
        },
    )

    async for message in query(prompt=prompt, options=options):
        if getattr(message, "type", None) == "system" and getattr(message, "subtype", None) == "mirror_error":
            print("SessionStore mirror failed", message)

        if getattr(message, "type", None) == "result":
            print(
                {
                    "subtype": getattr(message, "subtype", None),
                    "cost": getattr(message, "total_cost_usd", None),
                    "usage": getattr(message, "usage", None),
                    "structured": getattr(message, "structured_output", None),
                }
            )


asyncio.run(run_agent("Analyze this tenant workspace and return a structured summary"))

Python env overlays inherited variables. Prefer container secrets or gateway injection for authentication and proxy settings.

#Launch checklist

CheckAcceptance criterion
Process boundaryExplicit cwd, resources, turn limits and budget per session
RecoverySessionStore configured where required; mirror_error monitored
PersistenceSeparate retention policies for memory, artifacts and transcripts
IsolationEmpty setting sources, disabled auto memory, per-tenant config and egress implemented
ObservabilityOTEL traces/metrics/logs reach collector; sensitive logging off by default
CostsResult costs, usage and cache reads/writes archived by tenant
ToolsAllowlisted tools/MCP; search evaluated for large inventories
CachingDefault five minutes or one-hour override chosen for session intervals

#Official references

Support

Need help?

For setup, billing, or model issues, email us. Check the status page for uptime.

WeChat / QQ support is available at the bottom right.