Gateway operations
Operate apps and LLM gateways with authentication, routing, caching, discovery and spend limits.
There are two operational forms: administrator-hosted apps gateway for organizational access, and generic LLM gateways such as Passion8 forwarding Anthropic Messages from ANTHROPIC_BASE_URL.
The apps-gateway login, callback and public URL must be reachable by user devices within the controlled network, not only from the server itself.
See protocol checklist for forwarding, errors, discovery and caching; see provider guidance for user authentication variables.
#Official apps gateway
Administrators host it with the claude binary:
claude gateway --config gateway.yamlNo separate application framework is required. Configure identity, policy, database, telemetry and routing, with login behind VPN, private DNS or controlled proxy.
| Capability | Operational concern |
|---|---|
| OIDC | Issuer, client, callback and claims; verify device access to login/callback |
| Postgres | Backups, migrations, pools and monitoring for persistent state |
| Managed policy | Central model/tool/permission defaults instead of manual user settings |
| Telemetry | Requests, latency, errors, routes, spend and policy; no raw prompts by default |
| Routing | Route by user, team, model, cost or region |
| Spend limits | User/team/project budgets and explicit errors, not silent model substitution |
#gateway.yaml example
This operational outline illustrates configuration areas. Use the current gateway schema for exact field names.
server:
listen: "0.0.0.0:8443"
public_url: "https://claude-gateway.internal.example.com"
auth:
oidc:
issuer_url: "https://idp.internal.example.com"
client_id: "claude-apps-gateway"
client_secret_env: "CLAUDE_GATEWAY_OIDC_CLIENT_SECRET"
allowed_groups:
- "engineering"
- "support"
database:
postgres:
url_env: "CLAUDE_GATEWAY_DATABASE_URL"
policies:
managed:
enabled: true
default_policy: "engineering-default"
telemetry:
otlp:
endpoint: "https://otel.internal.example.com/v1/traces"
headers_env: "CLAUDE_GATEWAY_OTEL_HEADERS"
upstreams:
- name: "anthropic"
type: "anthropic"
base_url: "https://api.anthropic.com"
api_key_env: "ANTHROPIC_API_KEY"
- name: "passion8"
type: "anthropic"
base_url: "https://passion8.cc"
auth_token_env: "PASSION8_API_KEY"
routing:
default_upstream: "anthropic"
rules:
- group: "support"
upstream: "passion8"
models:
- "claude-sonnet-5-5"
spend_limits:
defaults:
user_daily_usd: 10
team_monthly_usd: 1000
on_exceeded: "deny"Put OIDC secrets, database URL, upstream keys and telemetry headers in environment variables or secret management, not plaintext gateway.yaml.
#Generic LLM gateway protocol
Code still sends Anthropic Messages through Passion8/custom gateways. Set the root in ANTHROPIC_BASE_URL and provide compatible paths.
| Item | Requirement |
|---|---|
| Base URL | Host root such as https://passion8.cc, usually without /v1 |
| Messages | POST /v1/messages required |
| Token count | Optional; absence causes estimates or reduced functionality |
| Streaming | SSE for stream:true without proxy buffering |
| Headers | Preserve version/beta headers rather than forcing old values |
| Errors | Preserve JSON status, type and message for retries/fallback |
Use extensible validation rather than a minimal fixed schema. Preserve these fields and explicitly forward or reject new beta fields:
| Field | Meaning |
|---|---|
| model, max_tokens, messages, system | Core Messages fields |
| stream | SSE control |
| tools, tool_choice | Built-in/MCP/deferred tools |
| thinking | Reasoning/effort |
| metadata, stop_sequences | Business metadata and termination |
| temperature, top_p, top_k | Sampling |
| context_management, output_config | Context and output control |
| cache_control | Nested system/message/tool blocks |
#System arrays and attribution
Preserve system array block order, types and nested fields, including the attribution block's position.
| Avoid | Risk |
|---|---|
| Flattening arrays | Breaks attribution, cache keys and controls |
| Prepending custom blocks | Attribution may enter the actual prompt and prefix |
| Sorting/deduplicating/rewriting | Cache misses and policy/capability errors |
| Dropping unknown fields | New beta features fail |
Prefer managed/provider policy. If rewriting is unavoidable, preserve attribution ordering and audit the transformation.
#Model discovery
Enable gateway models in the picker with:
export CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1Discovery request:
GET /v1/models?limit=1000Operational constraints:
| Item | Details |
|---|---|
| Timeout | About three seconds; slow responses, redirects or OIDC cause silent failure |
| Cache | Usually ~/.claude/cache/gateway-models.json |
| Manual defaults | ANTHROPIC_MODEL and DEFAULT_OPUS/SONNET/HAIKU/FABLE_MODEL variables |
| Capabilities | Listing does not establish search, 1M context, thinking or one-hour caching |
Return stable, lightweight, noninteractive model JSON; never redirect discovery to an IdP or login page.
#One-hour prompt caching
A client variable alone cannot guarantee one-hour caching. Verify:
| Condition | Check |
|---|---|
| Client | Enable one-hour TTL without FORCE_PROMPT_CACHING_5M or DISABLE_PROMPT_CACHING |
| Upstream | Actual routed provider/model supports it |
| Headers | No version/beta stripping or downgrade |
| Stable body | Preserve arrays, messages, tools, cache controls and model parameters |
| Usage | Preserve creation/read token fields |
Document and observe five-minute fallback if one hour is unsupported. Never fabricate hits through body rewriting.
#Spend limits
Consider budgets at gateway and provider levels:
| Layer | Recommendation |
|---|---|
| Gateway | Daily/monthly limits and concurrency by user/team/project/model/upstream |
| Provider | Hard caps against bugs or credential exposure |
| Response | Stable JSON 402/403/429 with limit, window and reset |
| Observability | Spend, denials, matched rules and fallback |
Do not silently substitute a cheaper, different model after a limit unless policy and UI explicitly explain it. Silent fallback complicates debugging, cost attribution and auditing.
#Server-managed settings limitations
Official settings depend on Anthropic identities and policy services. Do not assume they govern third-party providers or deliver custom gateway URLs/tokens.
| Scenario | Recommendation |
|---|---|
| Standard URL | MDM, endpoint/system settings or templates |
| Token delivery | apiKeyHelper, vault or short-lived credentials; no repository secrets |
| Models/budgets | Gateway policies and spend limits |
| Apps gateway | Its OIDC, Postgres, managed policy and telemetry |
#Passion8/custom gateway checklist
| Check | Details |
|---|---|
| Base URL | Root such as https://passion8.cc, not OpenAI /v1 |
| Auth | Usually ANTHROPIC_AUTH_TOKEN; avoid stale API-key/login conflicts |
| Private access | Login, callback and admin/health reachable on controlled network |
| Proxy | Unbuffered SSE, persistent connections and suitable idle timeout |
| Forwarding | Preserve headers, arrays, tools, caching, usage and errors |
| Audit | Request ID, user, model, tokens and policy; no full payload by default |
#curl self-checks
claude-sonnet-5-5 is a current official example ID. Verify Passion8 access first or substitute an enabled console model.
Set environment variables first:
export ANTHROPIC_BASE_URL="https://passion8.cc"
export ANTHROPIC_AUTH_TOKEN="sk-..."
export ANTHROPIC_MODEL="claude-sonnet-5-5"Messages JSON:
curl -sS "$ANTHROPIC_BASE_URL/v1/messages" \
-H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "'"$ANTHROPIC_MODEL"'",
"max_tokens": 64,
"messages": [
{ "role": "user", "content": "Reply only ok" }
]
}'SSE stream:
curl -N "$ANTHROPIC_BASE_URL/v1/messages" \
-H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "'"$ANTHROPIC_MODEL"'",
"max_tokens": 64,
"stream": true,
"messages": [
{ "role": "user", "content": "stream ok" }
]
}'Test discovery within the client-like three-second budget:
curl -sS --max-time 3 "$ANTHROPIC_BASE_URL/v1/models?limit=1000" \
-H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN" \
-H "anthropic-version: 2023-06-01"Optional token count:
curl -sS "$ANTHROPIC_BASE_URL/v1/messages/count_tokens" \
-H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "'"$ANTHROPIC_MODEL"'",
"messages": [
{ "role": "user", "content": "count tokens" }
]
}'Expected checks:
| Item | Passing condition |
|---|---|
| Messages | Anthropic JSON, not HTML/login/proxy pages |
| SSE | curl -N receives event/data incrementally without buffering |
| Models | JSON or explicit JSON 401/403 within three seconds |
| Count tokens | Real count when supported, explicit error otherwise |
| Spend cap | Stable JSON error with attributable telemetry |
#Official references
- Claude apps gateway deployment and operations
- Deploy Claude apps gateway on Google Cloud
- Run Claude Code through a gateway
#Related pages
Support
Need help?
For setup, billing, or model issues, email us. Check the status page for uptime.
WeChat / QQ support is available at the bottom right.

