Prompt caching
Automatic caching, five-minute and one-hour TTLs, invalidation, MCP/tool search, and usage fields.
Claude Code automatically uses prompt caching. You normally do not need to write cache_control yourself, but should understand which actions make the next turn slower or more expensive and how TTL choices differ.
See Commands and cache effects for command-by-command details. This page covers mechanics and diagnosis.
#Core concept
Every request resends system instructions, project context, history, tool results, and the new message. A cache hit lets the upstream reuse the identical prefix and process the appended content.
Matching starts at the beginning of the request and requires an identical prefix. A change inside that prefix affects reuse after that point.
#Context layers
| Layer | Contents | Common changes |
|---|---|---|
| System prompt | Core instructions, tools, output style | CLI upgrade, tool definitions, rebuilt style |
| Project context | CLAUDE.md, auto memory, unconditional rules | Startup, clear, compact |
| Conversation | User/assistant messages and tool results | Appended every turn |
Earlier changes affect more of the prefix. Appending conversation content is generally easier to reuse than rebuilding the system prompt.
Cache identity also involves:
- Model: switching models does not reuse the old model's cache.
- Effort: changing effort can invalidate reuse.
- Fast-mode headers: initial activation can require a new cache.
Advisor is a special case: its tool definition is placed after the cache breakpoint, so toggling it generally preserves the existing prefix.
#Five minutes versus one hour
| Item | Five-minute TTL | One-hour TTL |
|---|---|---|
| Default audience | API keys, cloud providers, third-party providers | Subscriptions within included usage |
| Official cache-write cost | Commonly 1.25 times base input price | Commonly twice base input price |
| Read cost | Current model-specific cache-read price | Same model-specific cache-read price |
| Suitable for | Continuous short-gap iteration | Large contexts with longer pauses |
| Enablement | Default without ENABLE_PROMPT_CACHING_1H | Explicit ENABLE_PROMPT_CACHING_1H=1 where applicable |
| Force | FORCE_PROMPT_CACHING_5M=1 | Not applicable |
A hit refreshes TTL; five minutes is an inactivity interval, not the total lifetime. Official pricing is not a promise of Passion8 billing.
Do not put one-hour caching into every API/provider template. Enable it for a workload that benefits from longer reuse. Included subscription requests can receive one-hour caching automatically; usage-credit requests follow their documented TTL rules.
#Actions that invalidate reuse
| Action | Reason |
|---|---|
| Switch model | Separate model cache |
| Change effort | Changes reasoning configuration/cache identity |
| Enable fast mode mid-session | Request header changes |
| Enter/leave plan under opusplan | Can switch Opus and Sonnet |
| Fallback activates | Another model is used |
| Connect/disconnect MCP | Definitions may enter the prefix |
| Toggle a plugin with MCP | Tool set changes |
| Deny an entire tool | Tool removed from context |
| Compact | Summary replaces conversation history |
| Upgrade Claude Code | System instructions or tools can change |
MCP effects depend on whether definitions are deferred. Supported tool search improves prefix stability. Custom gateways or unsupported provider/model combinations can load tools upfront, making connection changes more disruptive.
#Changes that do not immediately replace the prefix
| Action | Result |
|---|---|
| Edit project files | Appends a file-change notification |
| Edit root CLAUDE.md | Current session may retain loaded instructions |
| Change outputStyle | Loaded system prompt does not immediately change |
| Change permission mode | Usually does not alter the prompt |
| Invoke skill/command | Appends a message |
| /recap | Adds a summary without replacing history |
| /cd | Preserves conversation where possible; appends directory context |
| /advisor | Definition after the breakpoint generally preserves prefix |
| /btw | Side question stays outside main conversation |
| /goal | Goal state does not itself rebuild system instructions |
| /loop | Ordinary new turns; interval determines TTL warmth |
| /rewind | Returns to an earlier prefix that may remain cached |
| /statusline | Local script refresh, no model request |
| /usage and /cost | Inspection only |
| Enable OTel | Normally no prompt change |
New root instructions and output styles apply when context is reloaded through the relevant new-session/compaction flow.
#Compaction cost
Compaction makes a summarization request, which can often read the existing prefix because the summary instruction is appended to the history.
Afterward, short summary content replaces the old conversation, and the next turn creates a cache for that shorter history.
- Compact at a natural task boundary.
- Avoid waiting for automatic compaction in a critical step.
- Use rewind, not compact, to abandon an incorrect direction.
#Subagents and forks
Subagents have separate conversations and caches. Their first call is usually cold; subsequent turns warm their own prefix. Their documented default TTL can differ from the subscription main conversation.
The invocation and final result append to the main conversation rather than replacing its prefix.
Fork/branch copies the conversation more directly, including instructions, tools, and history. Within TTL, the first fork turn may reuse the parent's prefix. Worktrees primarily isolate files and Git branches; see Worktrees.
#Usage fields
| Field | Meaning |
|---|---|
| cache_creation_input_tokens | Tokens written this turn, billed at cache-write rates |
| cache_read_input_tokens | Tokens reused this turn, billed at current model/gateway read rates |
A higher read-to-write ratio suggests better reuse. Repeated creation means prefixes may be changing: check model, effort, style, MCP, tool denials, compaction, and CLI upgrades.
On Passion8 or another gateway, verify forwarding of anthropic-beta, tool schemas, output_config, context_management, cache fields, and usage. See Gateway and protocol.
#Disable or force TTL
# Disable all prompt caching
export DISABLE_PROMPT_CACHING=1
# Disable a model family
export DISABLE_PROMPT_CACHING_SONNET=1
export DISABLE_PROMPT_CACHING_OPUS=1
export DISABLE_PROMPT_CACHING_HAIKU=1
export DISABLE_PROMPT_CACHING_FABLE=1
# Optional one-hour TTL for API/provider requests
export ENABLE_PROMPT_CACHING_1H=1
# Force five minutes over one-hour settings
export FORCE_PROMPT_CACHING_5M=1Disable caching only for diagnosis, not ordinary use.
#Cache-friendly workflow
Choose model and effort first
Choose /model and /effort before starting the task.
Avoid repeatedly switching during the work.Keep stable information early
Store durable project conventions in CLAUDE.md and define task scope up front. Keep temporary errors, logs, and casual TODOs out of permanent memory.
Compact at natural breaks
Compact after a coherent unit of work. Use clear for an unrelated task.
#Official references
Support
Need help?
For setup, billing, or model issues, email us. Check the status page for uptime.
WeChat / QQ support is available at the bottom right.

