Claude Code

Agent SDK streaming and plugins

Streaming input, single messages, plugin loading, skill namespaces, session controls, permissions and cache effects.

Input modes and plugin loading are important runtime decisions. Input mode determines process lifetime, queued messages, image input, interruption and recovery. Plugins determine access to skills, agents, hooks and MCP servers.

Read the API reference and runtime patterns before choosing a product design here.

#Compare input modes

ModeBest fitLimitations
Streaming inputLong tasks, chat UIs, dashboards, images, queues and interruptionHandle generator errors, permission pauses and cancellation
Single messageOne-shot analysis, Lambda/jobs and noninteractive scriptsNo direct image attachments, dynamic queuing, realtime interruption or natural multiple turns

Official guidance recommends streaming input by default: a long-lived process continuously receives messages, runs tools, emits intermediate status and retains context in one session.

#Streaming input structure

TypeScript streaming input
import { query, type SDKUserMessage } from "@anthropic-ai/claude-agent-sdk";

async function* messages(): AsyncGenerator<SDKUserMessage> {
  yield {
    type: "user",
    message: {
      role: "user",
      content: "Analyze this codebase for security issues"
    },
    parent_tool_use_id: null
  };

  await new Promise((resolve) => setTimeout(resolve, 2000));

  yield {
    type: "user",
    message: {
      role: "user",
      content: "Now focus on auth and payment paths"
    },
    parent_tool_use_id: null
  };
}

for await (const message of query({
  prompt: messages(),
  options: {
    maxTurns: 10,
    allowedTools: ["Read", "Grep"]
  }
})) {
  if (message.type === "result") console.log(message);
}
Python streaming input
from claude_agent_sdk import ClaudeSDKClient, ClaudeAgentOptions
import asyncio

async def main():
    async def messages():
        yield {
            "type": "user",
            "message": {
                "role": "user",
                "content": "Analyze this codebase for security issues",
            },
        }
        await asyncio.sleep(2)
        yield {
            "type": "user",
            "message": {
                "role": "user",
                "content": "Now focus on auth and payment paths",
            },
        }

    async with ClaudeSDKClient(ClaudeAgentOptions(max_turns=10)) as client:
        await client.query(messages())
        async for message in client.receive_response():
            print(message)

asyncio.run(main())

A TypeScript generator error may appear as an aborted session. Python generator errors may appear only in debug logs and leave the session hanging. Catch file-read, image-encoding and network errors inside your generators.

#Single-message scenarios

One-shot query
import { query } from "@anthropic-ai/claude-agent-sdk";

for await (const message of query({
  prompt: "Explain the authentication flow",
  options: {
    maxTurns: 1,
    allowedTools: ["Read", "Grep"]
  }
})) {
  if (message.type === "result" && message.subtype === "success") {
    console.log(message.result);
  }
}

For error_max_turns or another error result, single-message query() may throw after yielding its final result. Wrap the entire async loop in try/catch rather than checking only success.

#Load SDK plugins

The SDK accepts local plugin paths only:

TypeScript plugins
import { query } from "@anthropic-ai/claude-agent-sdk";

for await (const message of query({
  prompt: "What custom capabilities do you have?",
  options: {
    plugins: [
      { type: "local", path: "./plugins/review" },
      { type: "local", path: "/opt/company/claude-plugins/security" }
    ],
    maxTurns: 3
  }
})) {
  if (message.type === "system" && message.subtype === "init") {
    console.log(message.plugins);
    console.log(message.skills);
    console.log(message.slash_commands);
  }
}
Python plugins
from claude_agent_sdk import query, ClaudeAgentOptions
import asyncio

async def main():
    async for message in query(
        prompt="What custom capabilities do you have?",
        options=ClaudeAgentOptions(
            plugins=[
                {"type": "local", "path": "./plugins/review"},
                {"type": "local", "path": "/opt/company/claude-plugins/security"},
            ],
            max_turns=3,
        ),
    ):
        print(message)

asyncio.run(main())

Point to the plugin root containing skills/, agents/, hooks/, commands/ or .claude-plugin/. CLI-installed plugins can be reused by passing their installed local path, usually under ~/.claude/plugins/.

#Plugin skills and namespaces

Plugin skills automatically receive a plugin-name prefix to avoid built-in command collisions.

ContentSDK representation
Skill security in review-kit/review-kit:security
Legacy commands/custom-command/review-kit:custom-command
Plugin agentAgents/capabilities in system init
Plugin MCP serverUsually mcp__server__tool names

Do not hard-code the UI's command inventory. Read skills, slash_commands, plugins and available tools from system/init and render those capabilities.

#Permissions and pauses

Handle three kinds of pause in streaming mode:

PauseHandling
Tool approvalShow tool, argument summary and risk; return allow/deny/defer
AskUserQuestionPresent a form or choices, then resume
Plugin hook blocksShow plugin name, reason and recovery guidance

Tools preapproved by allowedTools, permission rules or modes do not invoke canUseTool. Put mandatory per-call checks in hooks or gateway policy.

#Caching and cost

Change5m / 1h effect
Continuous streaming inputStable system, tools and settings favor warm cache reuse
Stateless single-message jobsEach resembles a cold start; limited five-minute reuse
Changed plugin listSkills, hooks, MCP and schemas change the prefix
Plugin MCP serversMany upfront tools increase cache creation substantially
Tool searchFewer schemas enter the prefix, improving cache stability
/compactRewrites history; the next turn may miss cache

The SDK uses prompt caching automatically, but one-hour TTL depends on model, provider, environment and gateway forwarding. With Passion8, observe whether cache_read_input_tokens grows consistently.

#Official references

Support

Need help?

For setup, billing, or model issues, email us. Check the status page for uptime.

WeChat / QQ support is available at the bottom right.