Claude Code

Models, effort, and Fast Mode

/model, /effort, opusplan, fallback models, Fast Mode, 1M context, provider capabilities, and caching.

Model configuration affects quality, speed, cost, and caching. With Passion8, verify the actual model ID, capability metadata, context window, effort, and Fast Mode support.

For advisor-assisted decisions, Agent View, goals, Ultraplan, and Ultrareview, see Advanced workflows.

#Model selection precedence

SourceBehaviorUse
/model <alias or id>Switch immediately; current versions can save a user defaultInteractive selection
claude --model <alias or id>This launch onlyDifferent models in separate terminals
ANTHROPIC_MODELSessions launched with that environmentShell profiles or scripts
settings.json modelPersistent defaultPersonal/team preferences
Managed settingsAdministrator policyOrganization restrictions

Resuming a session normally uses its saved model unless unavailable, excluded by policy, or explicitly overridden by the new launch's model selection.

Switching model changes cache identity. The next turn must process history for the new model; neither five-minute nor one-hour TTL transfers it across models.

#Common aliases

AliasMeaningNotes
defaultAccount/provider/organization defaultVaries by plan and provider
opusDefault Opus candidatePlanning, architecture, difficult review
sonnetDefault Sonnet candidateMost coding execution
haikuDefault Haiku candidateLightweight work; capabilities vary
fableDefault Fable candidate; checked catalog lists Fable 5.1Safety/bio content may trigger fallback
opusplanOpus for planning, Sonnet for executionMode transitions can switch model
opus[1m] / sonnet[1m]Request a 1M-context variantRequires model/account/provider support

Use the actual Passion8 ID if the gateway uses custom names. For model discovery in the picker, see Gateway and protocol.

#opusplan

PhaseModelUse
Plan modeOpusRequirements, architecture, risk
ExecutionSonnetEditing, testing, ordinary implementation

The tradeoff is split caches when the actual model changes.

  1. Plan at the beginning with opusplan.
  2. After approval, avoid repeatedly moving between planning and execution.
  3. For an occasional second opinion, consider advisor instead of switching the whole session.

#Effort

Effort controls adaptive reasoning depth. Higher settings can cost more and take longer; supported levels and defaults depend on the model.

EffortTypical use
lowShort, deterministic, latency-sensitive work
mediumModerate tasks with cost constraints
highBalanced coding work where supported
xhighComplex architecture, migration, debugging
maxDeep session work; avoid an unconditional global default
ultracodeClaude Code mode combining xhigh and more active dynamic workflows

Set at launch:

claude --effort high

Or in-session:

/effort xhigh
/effort auto

Or in settings:

settings.json
{
  "effortLevel": "high"
}

Effort changes can invalidate cached history. Choose it before a long task. An explicit CLAUDE_CODE_EFFORT_LEVEL environment value takes precedence over session effort controls.

#Fast Mode

Fast Mode trades higher cost for lower latency on supported Opus configurations.

/fast
ScenarioRecommendation
Speed required from the startEnable before the long task
Existing long sessionInitial activation may require reprocessing history
Not currently OpusActivation may switch model
Rate limitedOfficial service can fall back to standard speed; cache behavior depends on version

Fast Mode is a speed/billing setting; effort controls reasoning depth. They are not interchangeable.

#Advisor tool

The main model continues working while consulting another model on difficult decisions.

/advisor opus
/advisor off

At startup:

claude --advisor opus
AspectCache effect
Toggle advisorGenerally preserves main prefix
Advisor requestReads the transcript; advisor calls do not share their own cache as ordinary continued turns
Advisor answerAppended to main history and can participate in later caching

Advisor is a server-executed Anthropic tool. Availability through cloud providers or third-party gateways depends on actual support and passthrough.

#Fallback models

Configure a fallback chain for overload/unavailability:

claude --fallback-model sonnet,haiku

Or in settings:

settings.json
{
  "fallbackModel": ["sonnet", "haiku"]
}

A fallback turn uses another model and therefore another cache. It improves availability rather than reducing cache cost.

#1M context

The official configuration supports [1m] aliases/suffixes on compatible models. Availability depends on model, account, provider, and gateway.

CheckReason
Actual 1M model routeAn alias alone cannot create upstream support
Context/beta field passthroughStripping fields can disable the requested capability
Compaction thresholdA smaller gateway window may require CLAUDE_CODE_AUTO_COMPACT_WINDOW

#Third-party capability declarations

Claude Code infers effort, thinking, tool search, and context capabilities from model IDs. Custom deployment names may need provider-specific declarations.

export ANTHROPIC_DEFAULT_OPUS_MODEL="my-opus-deployment"
export ANTHROPIC_DEFAULT_OPUS_MODEL_SUPPORTED_CAPABILITIES="effort,xhigh_effort,max_effort,thinking,adaptive_thinking,interleaved_thinking"

The official capability variables apply to supported provider configurations such as Bedrock, Vertex, Foundry, and Mantle. Do not assume the same behavior on every ANTHROPIC_BASE_URL gateway; full header/body forwarding remains essential.

#Passion8 recommendations

NeedRecommendation
Stable codingAvailable Sonnet/Opus ID; avoid frequent switching
Deep planningTry opusplan or supported advisor separately
Long contextVerify model window and gateway fields before [1m]
Cost controlChoose a supported moderate effort; compact at natural breaks
SpeedDecide on Fast Mode before a long conversation
Gateway model pickerRequires /v1/models and CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1

#Official references

Support

Need help?

For setup, billing, or model issues, email us. Check the status page for uptime.

WeChat / QQ support is available at the bottom right.