Gateways and protocols
Passion8 endpoints, authentication, forwarding, model discovery, caching, tool search and common 400/401 failures.
ANTHROPIC_BASE_URL connects Code to an LLM gateway. The CLI sends Anthropic Messages while Passion8 handles forwarding, authentication, routing, billing and observability.
Code uses https://passion8.cc without /v1. Codex and OpenAI-compatible clients use https://passion8.cc/v1.
For client configuration, start here and with provider authentication guidance. Implementers should use the protocol checklist.
#Minimal setup
export ANTHROPIC_BASE_URL="https://passion8.cc"
export ANTHROPIC_AUTH_TOKEN="sk-YOUR_PASSION8_API_KEY"
claude -p "Reply only ok"Windows PowerShell:
$env:ANTHROPIC_BASE_URL = "https://passion8.cc"
$env:ANTHROPIC_AUTH_TOKEN = "sk-YOUR_PASSION8_API_KEY"
claude -p "Reply only ok"Store keys in user settings or local environment, never project Git history:
{
"env": {
"ANTHROPIC_BASE_URL": "https://passion8.cc",
"ANTHROPIC_AUTH_TOKEN": "sk-YOUR_PASSION8_API_KEY"
}
}#Request format
Code uses Anthropic Messages for ANTHROPIC_BASE_URL:
| Item | Details |
|---|---|
| Main endpoint | /v1/messages |
| Optional endpoint | /v1/messages/count_tokens |
| Authentication | Bearer authorization or x-api-key |
| Streaming | SSE forwarding required |
| Models | Optional /v1/models?limit=1000 |
Provider dialect conversion belongs to the gateway. Do not use the OpenAI /v1 root for Code.
#Required forwarding
Avoid fixed allowlists that strip new fields. Client capabilities and beta headers evolve.
| Content | Importance |
|---|---|
| anthropic-version | Upstream version |
| anthropic-beta | Search, context and beta tool capabilities |
| system array ordering | Attribution, prompt and cache key |
| tools/schemas | MCP, built-ins and deferred loading |
thinking | adaptive reasoning |
output_config | effort、structured output、task budget |
| context_management | Automatic context controls |
| Error body | Retry and capability fallback decisions |
Auditing should inspect without rewriting bodies. Header stripping or flattening system arrays changes capabilities and caching.
#Attribution blocks and caching
Code prepends attribution; the official endpoint strips it when correctly positioned to protect first-party caching.
Custom gateway considerations:
| Behavior | Result |
|---|---|
| Preserve arrays and first attribution block | Most reliable |
| Prepend custom system block | Attribution may enter prompt/cache key |
| Flatten array | Can break stripping/caching |
| Required system rewriting | Consider CLAUDE_CODE_ATTRIBUTION_HEADER=0 |
From v2.1.181, attribution is more stable within conversations on custom base URLs, helping full-body caches. Older deployments may need the attribution override.
#Prompt caching through Passion8
Check three layers:
| Layer | Check |
|---|---|
| Client | Frequent model/effort/fast/compact changes or upgrades |
| Gateway | Cache controls, headers, tools and usage preserved |
| Model | Caching, search, one-hour TTL and context support |
Usage fields:
| Field | Meaning |
|---|---|
| cache_creation_input_tokens | Tokens written this turn |
| cache_read_input_tokens | Tokens read this turn |
A higher read/write ratio suggests better reuse. Repeated high writes warrant checking model, effort, fast mode, MCP schemas, plugins, tool denies and compaction.
#Model discovery
Implementing /v1/models lets gateway models appear in the picker.
Enable:
export CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1Notes:
| Condition | Explanation |
|---|---|
| ANTHROPIC_BASE_URL only | Higher-priority cloud-provider variables bypass this discovery |
| Short timeout | Slow or redirected discovery silently fails |
| Cached results | ~/.claude/cache/gateway-models.json |
| Custom capabilities | Picker presence does not prove effort, search or 1M context |
#Common errors
| Symptom | Possible cause | Fix |
|---|---|---|
| 401 | Wrong key/header or stale login | Verify token; log out if needed |
| ConnectionRefused | Endpoint, VPN or firewall | Test gateway with curl |
| 200 but malformed | HTML login/proxy response | Fix routing to return API JSON/SSE |
| 400 context_management | Unsupported upstream | Route to support or temporarily disable experimental betas |
| 400 thinking/adaptive | Unsupported reasoning | Upgrade upstream or use documented capability controls |
| No model list | Missing endpoint or disabled discovery | Enable discovery or set an actual model ID |
| Remote Control unavailable | Custom credential/base URL | See remote access |
| Official account features missing | Different provider/account path | See availability |
| Inaccurate context estimates | Missing count_tokens | Implement counting or accept estimates |
#Enterprise configuration
Teams can use:
| Capability | Purpose |
|---|---|
| Managed settings | Endpoint, models, permissions and hooks |
| Server-managed | Organizational policy for eligible cloud/web sessions |
| apiKeyHelper | Dynamic gateway tokens |
| Custom headers | Tenant, routing and audit metadata |
| OpenTelemetry | Usage, tools, hooks and costs |
| Gateway limits | Per-user daily/weekly/monthly caps |
Helper output is cached. CLAUDE_CODE_API_KEY_HELPER_TTL_MS controls the documented helper cache duration.
See operations and apps-gateway deployment for OIDC, Postgres, budgets, discovery, forwarding and TTL boundaries.
#Passion8 checklist
- Use the root URL without /v1.
- Use ANTHROPIC_AUTH_TOKEN and avoid conflicting credentials.
- Choose model, effort and fast mode before long tasks.
- Verify search/deferred tools for large MCP inventories.
- Diagnose zero reads with stable model/effort and no MCP.
- Verify the actual model ID before enabling discovery.
- For unavailable Remote Control or dictation, test the documented official-account boundary separately.
#Official references
- Run Claude Code through a gateway
- Gateway protocol reference
- Connect Claude Code to an LLM gateway
- Other LLM gateways
- Prompt caching
- Environment variables
- Feature availability
#Related pages
Support
Need help?
For setup, billing, or model issues, email us. Check the status page for uptime.
WeChat / QQ support is available at the bottom right.

