Claude Code

Cost optimization

Control Claude Code costs through context, models, caching, MCP, hooks, subagents and Passion8 gateway usage records.

Costs often grow because context accumulates, task boundaries become unclear and caches miss repeatedly. Control the workflow first, then examine cache fields.

Start a new task with /clear, plan a complex task with /plan, and use /compact at natural breaks in long sessions. Do not carry a morning installation error into an evening deployment task.

#Sources of cost

SourceWhy it growsControl
Input tokensHistory, files, logs and tool output enter context/clear, targeted @file references, shorter logs
Output tokensLong explanations, repeated summaries and unbounded code generationRequest short conclusions, tables, diffs or a plan
Tool outputLarge Bash, MCP, browser and database resultsPaginate, filter and limit output
Cache writesInitial processing of a stable prefix costs moreKeep stable content unchanged
Cache missesChanges to models, effort or MCP definitionsChoose settings early and avoid unnecessary switches

See context window for context composition and tool reference for output and permission behavior.

1

Start a new task

/clear

State the objective, scope and boundaries in one sentence.

2

Plan complex tasks first

Read the relevant files without editing. List the smallest change and verification commands.
3

Limit execution output

Implement the plan. Report only the changed files, verification performed and remaining risks.
4

Clean up at the end

After completion, use /compact, or use /clear when moving to another task.

#/clear, /compact and /context

CommandWhen to use itCost effect
/clearNew task, new module or irrelevant previous contextRemoves unrelated history from future inputs
/compactA long task must continueOne summarization request, followed by shorter context
/contextYou do not know what consumes tokensInspect usage before clearing
/rewindDiscard an incorrect directionReturns to an earlier prefix; often more suitable than compaction for undo

#Local UI does not always use tokens

FeatureCost
/statusline script refreshLocal execution; no model request or tokens
/usage, /costView usage without changing the prompt
/themeTerminal theme; no change to model requests
Change outputStyleChanges the system prompt when applied; the next turn costs more
OTel metrics/logsMainly local and observability-system costs, not additional model tokens

Use the status line to monitor context and cost instead of asking Claude to summarize costs every few turns.

#Models and effort

Model and effort participate in cache matching. Switching during a session can make the next turn miss its earlier cache.

  • Select the model with /model at the beginning.
  • Select effort with /effort at the beginning.
  • Avoid repeatedly switching Sonnet/Opus in a long session.
  • Run stronger-model review separately near the end of a task.

#Cache strategy

See prompt caching for the mechanism.

ScenarioRecommendation
Repeated changes within five minutesDefault caching is usually sufficient
Large context with pauses of 10–50 minutesAPI/provider users can consider ENABLE_PROMPT_CACHING_1H=1
Claude subscription within plan limitsClaude Code requests one-hour TTL automatically
Passion8 or a custom gatewayCheck whether usage includes cache reads and writes
Cache fields remain zeroCheck prompt length, gateway, model, effort and MCP

API price multipliers:

TypeMultiplier
Five-minute cache writeApproximately 1.25× input price
One-hour cache writeApproximately 2× input price
Cache readApproximately 0.1× input price

#Memory and rules

CLAUDE.md enters context at startup. It saves repeated explanation, but an excessively long file raises the fixed input cost.

  • Keep root CLAUDE.md within about 200 lines.
  • Put long workflows in skills or custom commands.
  • Place module-specific rules in .claude/rules/ using paths.
  • Keep temporary errors and today's TODO list out of permanent memory.

#MCP cost

MCP has two context costs:

  1. Tool definitions.
  2. Tool return values.

Reduce them by:

  • Connecting only the servers needed for the task.
  • Using tool search for large toolsets when supported by the provider/gateway.
  • Applying limits to database queries.
  • Extracting only relevant DOM or text from browser tools.
  • Filtering tool output before raising MAX_MCP_OUTPUT_TOKENS.

Tool search may be disabled by default with a custom ANTHROPIC_BASE_URL. Passion8 support depends on forwarding the required fields.

#Subagent cost

Subagents isolate large investigations: the main session receives their final findings rather than every file read. They still incur their own costs.

ScenarioGuidance
Large-repository investigation, approach comparison or parallel reviewSubagents are useful
Small edit to one small fileUsually unnecessary
Repeated subagent turnsAccount for their own caches and five-minute TTL
/fork copies current contextMay reuse the parent's cache, but also carries more context
Parallel edits in one repositoryIsolate changes with worktrees

#Hooks reduce rework

A hook may not reduce the current turn's tokens, but can prevent repeated repairs:

HookValue
Lightweight PostToolUse formattingAvoid asking the model to fix formatting repeatedly
PreToolUse command checksPrevent costly rollback after unsafe operations
Stop summary recordingLeave an audit record after long tasks
UserPromptSubmit sensitive-data checksKeep keys and private data out of context

Keep hooks short. Running a complete build after every edit can slow the workflow; run it at task completion where appropriate.

#Patterns to avoid

  • Mixing installation, coding, casual discussion and deployment in one session.
  • Pasting entire logs rather than relevant errors.
  • Asking to inspect a whole project without a goal.
  • Changing scope repeatedly during a task.
  • Connecting many MCP tools that the task does not use.
  • Storing keys, temporary errors and TODOs in CLAUDE.md.

Use /clear and restate a narrower task when these patterns accumulate.

#Official references

Support

Need help?

For setup, billing, or model issues, email us. Check the status page for uptime.

WeChat / QQ support is available at the bottom right.