Agent SDK streaming and plugins
Streaming input, single messages, plugin loading, skill namespaces, session controls, permissions and cache effects.
Input modes and plugin loading are important runtime decisions. Input mode determines process lifetime, queued messages, image input, interruption and recovery. Plugins determine access to skills, agents, hooks and MCP servers.
Read the API reference and runtime patterns before choosing a product design here.
#Compare input modes
| Mode | Best fit | Limitations |
|---|---|---|
| Streaming input | Long tasks, chat UIs, dashboards, images, queues and interruption | Handle generator errors, permission pauses and cancellation |
| Single message | One-shot analysis, Lambda/jobs and noninteractive scripts | No direct image attachments, dynamic queuing, realtime interruption or natural multiple turns |
Official guidance recommends streaming input by default: a long-lived process continuously receives messages, runs tools, emits intermediate status and retains context in one session.
#Streaming input structure
import { query, type SDKUserMessage } from "@anthropic-ai/claude-agent-sdk";
async function* messages(): AsyncGenerator<SDKUserMessage> {
yield {
type: "user",
message: {
role: "user",
content: "Analyze this codebase for security issues"
},
parent_tool_use_id: null
};
await new Promise((resolve) => setTimeout(resolve, 2000));
yield {
type: "user",
message: {
role: "user",
content: "Now focus on auth and payment paths"
},
parent_tool_use_id: null
};
}
for await (const message of query({
prompt: messages(),
options: {
maxTurns: 10,
allowedTools: ["Read", "Grep"]
}
})) {
if (message.type === "result") console.log(message);
}from claude_agent_sdk import ClaudeSDKClient, ClaudeAgentOptions
import asyncio
async def main():
async def messages():
yield {
"type": "user",
"message": {
"role": "user",
"content": "Analyze this codebase for security issues",
},
}
await asyncio.sleep(2)
yield {
"type": "user",
"message": {
"role": "user",
"content": "Now focus on auth and payment paths",
},
}
async with ClaudeSDKClient(ClaudeAgentOptions(max_turns=10)) as client:
await client.query(messages())
async for message in client.receive_response():
print(message)
asyncio.run(main())A TypeScript generator error may appear as an aborted session. Python generator errors may appear only in debug logs and leave the session hanging. Catch file-read, image-encoding and network errors inside your generators.
#Single-message scenarios
import { query } from "@anthropic-ai/claude-agent-sdk";
for await (const message of query({
prompt: "Explain the authentication flow",
options: {
maxTurns: 1,
allowedTools: ["Read", "Grep"]
}
})) {
if (message.type === "result" && message.subtype === "success") {
console.log(message.result);
}
}For error_max_turns or another error result, single-message query() may throw after yielding its final result. Wrap the entire async loop in try/catch rather than checking only success.
#Load SDK plugins
The SDK accepts local plugin paths only:
import { query } from "@anthropic-ai/claude-agent-sdk";
for await (const message of query({
prompt: "What custom capabilities do you have?",
options: {
plugins: [
{ type: "local", path: "./plugins/review" },
{ type: "local", path: "/opt/company/claude-plugins/security" }
],
maxTurns: 3
}
})) {
if (message.type === "system" && message.subtype === "init") {
console.log(message.plugins);
console.log(message.skills);
console.log(message.slash_commands);
}
}from claude_agent_sdk import query, ClaudeAgentOptions
import asyncio
async def main():
async for message in query(
prompt="What custom capabilities do you have?",
options=ClaudeAgentOptions(
plugins=[
{"type": "local", "path": "./plugins/review"},
{"type": "local", "path": "/opt/company/claude-plugins/security"},
],
max_turns=3,
),
):
print(message)
asyncio.run(main())Point to the plugin root containing skills/, agents/, hooks/, commands/ or .claude-plugin/. CLI-installed plugins can be reused by passing their installed local path, usually under ~/.claude/plugins/.
#Plugin skills and namespaces
Plugin skills automatically receive a plugin-name prefix to avoid built-in command collisions.
| Content | SDK representation |
|---|---|
| Skill security in review-kit | /review-kit:security |
| Legacy commands/custom-command | /review-kit:custom-command |
| Plugin agent | Agents/capabilities in system init |
| Plugin MCP server | Usually mcp__server__tool names |
Do not hard-code the UI's command inventory. Read skills, slash_commands, plugins and available tools from system/init and render those capabilities.
#Permissions and pauses
Handle three kinds of pause in streaming mode:
| Pause | Handling |
|---|---|
| Tool approval | Show tool, argument summary and risk; return allow/deny/defer |
| AskUserQuestion | Present a form or choices, then resume |
| Plugin hook blocks | Show plugin name, reason and recovery guidance |
Tools preapproved by allowedTools, permission rules or modes do not invoke canUseTool. Put mandatory per-call checks in hooks or gateway policy.
#Caching and cost
| Change | 5m / 1h effect |
|---|---|
| Continuous streaming input | Stable system, tools and settings favor warm cache reuse |
| Stateless single-message jobs | Each resembles a cold start; limited five-minute reuse |
| Changed plugin list | Skills, hooks, MCP and schemas change the prefix |
| Plugin MCP servers | Many upfront tools increase cache creation substantially |
| Tool search | Fewer schemas enter the prefix, improving cache stability |
/compact | Rewrites history; the next turn may miss cache |
The SDK uses prompt caching automatically, but one-hour TTL depends on model, provider, environment and gateway forwarding. With Passion8, observe whether cache_read_input_tokens grows consistently.
#Official references
#Related pages
Support
Need help?
For setup, billing, or model issues, email us. Check the status page for uptime.
WeChat / QQ support is available at the bottom right.

