# Claude Code Cost optimization

> Control Claude Code costs through context, models, caching, MCP, hooks, subagents and Passion8 gateway usage records.

URL: https://docs.passion8.cc/en/docs/claude-code/cost-saving
Language: en
Publisher: Passion8

Costs often grow because context accumulates, task boundaries become unclear and caches miss repeatedly. Control the workflow first, then examine cache fields.




Start a new task with `/clear`, plan a complex task with `/plan`, and use `/compact` at natural breaks in long sessions. Do not carry a morning installation error into an evening deployment task.




## Sources of cost

| Source | Why it grows | Control |
| --- | --- | --- |
| Input tokens | History, files, logs and tool output enter context | `/clear`, targeted `@file` references, shorter logs |
| Output tokens | Long explanations, repeated summaries and unbounded code generation | Request short conclusions, tables, diffs or a plan |
| Tool output | Large Bash, MCP, browser and database results | Paginate, filter and limit output |
| Cache writes | Initial processing of a stable prefix costs more | Keep stable content unchanged |
| Cache misses | Changes to models, effort or MCP definitions | Choose settings early and avoid unnecessary switches |

See [context window](https://docs.passion8.cc/en/docs/claude-code/context) for context composition and [tool reference](https://docs.passion8.cc/en/docs/claude-code/tools) for output and permission behavior.

## Recommended workflow





### Start a new task


```text
/clear
```

State the objective, scope and boundaries in one sentence.



### Plan complex tasks first


```text
Read the relevant files without editing. List the smallest change and verification commands.
```



### Limit execution output


```text
Implement the plan. Report only the changed files, verification performed and remaining risks.
```



### Clean up at the end

After completion, use `/compact`, or use `/clear` when moving to another task.






## `/clear`, `/compact` and `/context`

| Command | When to use it | Cost effect |
| --- | --- | --- |
| `/clear` | New task, new module or irrelevant previous context | Removes unrelated history from future inputs |
| `/compact` | A long task must continue | One summarization request, followed by shorter context |
| `/context` | You do not know what consumes tokens | Inspect usage before clearing |
| `/rewind` | Discard an incorrect direction | Returns to an earlier prefix; often more suitable than compaction for undo |

## Local UI does not always use tokens

| Feature | Cost |
| --- | --- |
| `/statusline` script refresh | Local execution; no model request or tokens |
| `/usage`, `/cost` | View usage without changing the prompt |
| `/theme` | Terminal theme; no change to model requests |
| Change `outputStyle` | Changes the system prompt when applied; the next turn costs more |
| OTel metrics/logs | Mainly local and observability-system costs, not additional model tokens |

Use the [status line](https://docs.passion8.cc/en/docs/claude-code/statusline) to monitor context and cost instead of asking Claude to summarize costs every few turns.

## Models and effort

Model and effort participate in cache matching. Switching during a session can make the next turn miss its earlier cache.

- Select the model with `/model` at the beginning.
- Select effort with `/effort` at the beginning.
- Avoid repeatedly switching Sonnet/Opus in a long session.
- Run stronger-model review separately near the end of a task.

## Cache strategy

See [prompt caching](https://docs.passion8.cc/en/docs/claude-code/prompt-caching) for the mechanism.

| Scenario | Recommendation |
| --- | --- |
| Repeated changes within five minutes | Default caching is usually sufficient |
| Large context with pauses of 10–50 minutes | API/provider users can consider `ENABLE_PROMPT_CACHING_1H=1` |
| Claude subscription within plan limits | Claude Code requests one-hour TTL automatically |
| Passion8 or a custom gateway | Check whether usage includes cache reads and writes |
| Cache fields remain zero | Check prompt length, gateway, model, effort and MCP |

API price multipliers:

| Type | Multiplier |
| --- | --- |
| Five-minute cache write | Approximately 1.25× input price |
| One-hour cache write | Approximately 2× input price |
| Cache read | Approximately 0.1× input price |

## Memory and rules

`CLAUDE.md` enters context at startup. It saves repeated explanation, but an excessively long file raises the fixed input cost.

- Keep root `CLAUDE.md` within about 200 lines.
- Put long workflows in skills or custom commands.
- Place module-specific rules in `.claude/rules/` using `paths`.
- Keep temporary errors and today's TODO list out of permanent memory.

## MCP cost

MCP has two context costs:

1. Tool definitions.
2. Tool return values.

Reduce them by:

- Connecting only the servers needed for the task.
- Using tool search for large toolsets when supported by the provider/gateway.
- Applying limits to database queries.
- Extracting only relevant DOM or text from browser tools.
- Filtering tool output before raising `MAX_MCP_OUTPUT_TOKENS`.

Tool search may be disabled by default with a custom `ANTHROPIC_BASE_URL`. Passion8 support depends on forwarding the required fields.

## Subagent cost

[Subagents](https://docs.passion8.cc/en/docs/claude-code/subagents) isolate large investigations: the main session receives their final findings rather than every file read. They still incur their own costs.

| Scenario | Guidance |
| --- | --- |
| Large-repository investigation, approach comparison or parallel review | Subagents are useful |
| Small edit to one small file | Usually unnecessary |
| Repeated subagent turns | Account for their own caches and five-minute TTL |
| `/fork` copies current context | May reuse the parent's cache, but also carries more context |
| Parallel edits in one repository | Isolate changes with [worktrees](https://docs.passion8.cc/en/docs/claude-code/worktrees) |

## Hooks reduce rework

A hook may not reduce the current turn's tokens, but can prevent repeated repairs:

| Hook | Value |
| --- | --- |
| Lightweight `PostToolUse` formatting | Avoid asking the model to fix formatting repeatedly |
| `PreToolUse` command checks | Prevent costly rollback after unsafe operations |
| `Stop` summary recording | Leave an audit record after long tasks |
| `UserPromptSubmit` sensitive-data checks | Keep keys and private data out of context |

Keep hooks short. Running a complete build after every edit can slow the workflow; run it at task completion where appropriate.

## Patterns to avoid

- Mixing installation, coding, casual discussion and deployment in one session.
- Pasting entire logs rather than relevant errors.
- Asking to inspect a whole project without a goal.
- Changing scope repeatedly during a task.
- Connecting many MCP tools that the task does not use.
- Storing keys, temporary errors and TODOs in `CLAUDE.md`.

Use `/clear` and restate a narrower task when these patterns accumulate.

## Official references

- [Manage costs effectively](https://code.claude.com/docs/en/costs.md)
- [How Claude Code uses prompt caching](https://code.claude.com/docs/en/prompt-caching.md)
- [Explore the context window](https://code.claude.com/docs/en/context-window.md)
- [Claude API prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md)
- [Status line](https://code.claude.com/docs/en/statusline.md)
- [Monitoring](https://code.claude.com/docs/en/monitoring-usage.md)
