Claude Code

Prompt caching

Automatic caching, five-minute and one-hour TTLs, invalidation, MCP/tool search, and usage fields.

Claude Code automatically uses prompt caching. You normally do not need to write cache_control yourself, but should understand which actions make the next turn slower or more expensive and how TTL choices differ.

See Commands and cache effects for command-by-command details. This page covers mechanics and diagnosis.

#Core concept

Every request resends system instructions, project context, history, tool results, and the new message. A cache hit lets the upstream reuse the identical prefix and process the appended content.

Matching starts at the beginning of the request and requires an identical prefix. A change inside that prefix affects reuse after that point.

#Context layers

LayerContentsCommon changes
System promptCore instructions, tools, output styleCLI upgrade, tool definitions, rebuilt style
Project contextCLAUDE.md, auto memory, unconditional rulesStartup, clear, compact
ConversationUser/assistant messages and tool resultsAppended every turn

Earlier changes affect more of the prefix. Appending conversation content is generally easier to reuse than rebuilding the system prompt.

Cache identity also involves:

  • Model: switching models does not reuse the old model's cache.
  • Effort: changing effort can invalidate reuse.
  • Fast-mode headers: initial activation can require a new cache.

Advisor is a special case: its tool definition is placed after the cache breakpoint, so toggling it generally preserves the existing prefix.

#Five minutes versus one hour

ItemFive-minute TTLOne-hour TTL
Default audienceAPI keys, cloud providers, third-party providersSubscriptions within included usage
Official cache-write costCommonly 1.25 times base input priceCommonly twice base input price
Read costCurrent model-specific cache-read priceSame model-specific cache-read price
Suitable forContinuous short-gap iterationLarge contexts with longer pauses
EnablementDefault without ENABLE_PROMPT_CACHING_1HExplicit ENABLE_PROMPT_CACHING_1H=1 where applicable
ForceFORCE_PROMPT_CACHING_5M=1Not applicable

A hit refreshes TTL; five minutes is an inactivity interval, not the total lifetime. Official pricing is not a promise of Passion8 billing.

Do not put one-hour caching into every API/provider template. Enable it for a workload that benefits from longer reuse. Included subscription requests can receive one-hour caching automatically; usage-credit requests follow their documented TTL rules.

#Actions that invalidate reuse

ActionReason
Switch modelSeparate model cache
Change effortChanges reasoning configuration/cache identity
Enable fast mode mid-sessionRequest header changes
Enter/leave plan under opusplanCan switch Opus and Sonnet
Fallback activatesAnother model is used
Connect/disconnect MCPDefinitions may enter the prefix
Toggle a plugin with MCPTool set changes
Deny an entire toolTool removed from context
CompactSummary replaces conversation history
Upgrade Claude CodeSystem instructions or tools can change

MCP effects depend on whether definitions are deferred. Supported tool search improves prefix stability. Custom gateways or unsupported provider/model combinations can load tools upfront, making connection changes more disruptive.

#Changes that do not immediately replace the prefix

ActionResult
Edit project filesAppends a file-change notification
Edit root CLAUDE.mdCurrent session may retain loaded instructions
Change outputStyleLoaded system prompt does not immediately change
Change permission modeUsually does not alter the prompt
Invoke skill/commandAppends a message
/recapAdds a summary without replacing history
/cdPreserves conversation where possible; appends directory context
/advisorDefinition after the breakpoint generally preserves prefix
/btwSide question stays outside main conversation
/goalGoal state does not itself rebuild system instructions
/loopOrdinary new turns; interval determines TTL warmth
/rewindReturns to an earlier prefix that may remain cached
/statuslineLocal script refresh, no model request
/usage and /costInspection only
Enable OTelNormally no prompt change

New root instructions and output styles apply when context is reloaded through the relevant new-session/compaction flow.

#Compaction cost

Compaction makes a summarization request, which can often read the existing prefix because the summary instruction is appended to the history.

Afterward, short summary content replaces the old conversation, and the next turn creates a cache for that shorter history.

  • Compact at a natural task boundary.
  • Avoid waiting for automatic compaction in a critical step.
  • Use rewind, not compact, to abandon an incorrect direction.

#Subagents and forks

Subagents have separate conversations and caches. Their first call is usually cold; subsequent turns warm their own prefix. Their documented default TTL can differ from the subscription main conversation.

The invocation and final result append to the main conversation rather than replacing its prefix.

Fork/branch copies the conversation more directly, including instructions, tools, and history. Within TTL, the first fork turn may reuse the parent's prefix. Worktrees primarily isolate files and Git branches; see Worktrees.

#Usage fields

FieldMeaning
cache_creation_input_tokensTokens written this turn, billed at cache-write rates
cache_read_input_tokensTokens reused this turn, billed at current model/gateway read rates

A higher read-to-write ratio suggests better reuse. Repeated creation means prefixes may be changing: check model, effort, style, MCP, tool denials, compaction, and CLI upgrades.

On Passion8 or another gateway, verify forwarding of anthropic-beta, tool schemas, output_config, context_management, cache fields, and usage. See Gateway and protocol.

#Disable or force TTL

# Disable all prompt caching
export DISABLE_PROMPT_CACHING=1

# Disable a model family
export DISABLE_PROMPT_CACHING_SONNET=1
export DISABLE_PROMPT_CACHING_OPUS=1
export DISABLE_PROMPT_CACHING_HAIKU=1
export DISABLE_PROMPT_CACHING_FABLE=1

# Optional one-hour TTL for API/provider requests
export ENABLE_PROMPT_CACHING_1H=1

# Force five minutes over one-hour settings
export FORCE_PROMPT_CACHING_5M=1

Disable caching only for diagnosis, not ordinary use.

#Cache-friendly workflow

1

Choose model and effort first

Choose /model and /effort before starting the task.
Avoid repeatedly switching during the work.
2

Keep stable information early

Store durable project conventions in CLAUDE.md and define task scope up front. Keep temporary errors, logs, and casual TODOs out of permanent memory.

3

Compact at natural breaks

Compact after a coherent unit of work. Use clear for an unrelated task.

#Official references

Support

Need help?

For setup, billing, or model issues, email us. Check the status page for uptime.

WeChat / QQ support is available at the bottom right.