Claude Code

Gateway operations

Operate apps and LLM gateways with authentication, routing, caching, discovery and spend limits.

There are two operational forms: administrator-hosted apps gateway for organizational access, and generic LLM gateways such as Passion8 forwarding Anthropic Messages from ANTHROPIC_BASE_URL.

The apps-gateway login, callback and public URL must be reachable by user devices within the controlled network, not only from the server itself.

See protocol checklist for forwarding, errors, discovery and caching; see provider guidance for user authentication variables.

#Official apps gateway

Administrators host it with the claude binary:

claude gateway --config gateway.yaml

No separate application framework is required. Configure identity, policy, database, telemetry and routing, with login behind VPN, private DNS or controlled proxy.

CapabilityOperational concern
OIDCIssuer, client, callback and claims; verify device access to login/callback
PostgresBackups, migrations, pools and monitoring for persistent state
Managed policyCentral model/tool/permission defaults instead of manual user settings
TelemetryRequests, latency, errors, routes, spend and policy; no raw prompts by default
RoutingRoute by user, team, model, cost or region
Spend limitsUser/team/project budgets and explicit errors, not silent model substitution

#gateway.yaml example

This operational outline illustrates configuration areas. Use the current gateway schema for exact field names.

server:
  listen: "0.0.0.0:8443"
  public_url: "https://claude-gateway.internal.example.com"

auth:
  oidc:
    issuer_url: "https://idp.internal.example.com"
    client_id: "claude-apps-gateway"
    client_secret_env: "CLAUDE_GATEWAY_OIDC_CLIENT_SECRET"
    allowed_groups:
      - "engineering"
      - "support"

database:
  postgres:
    url_env: "CLAUDE_GATEWAY_DATABASE_URL"

policies:
  managed:
    enabled: true
    default_policy: "engineering-default"

telemetry:
  otlp:
    endpoint: "https://otel.internal.example.com/v1/traces"
    headers_env: "CLAUDE_GATEWAY_OTEL_HEADERS"

upstreams:
  - name: "anthropic"
    type: "anthropic"
    base_url: "https://api.anthropic.com"
    api_key_env: "ANTHROPIC_API_KEY"
  - name: "passion8"
    type: "anthropic"
    base_url: "https://passion8.cc"
    auth_token_env: "PASSION8_API_KEY"

routing:
  default_upstream: "anthropic"
  rules:
    - group: "support"
      upstream: "passion8"
      models:
        - "claude-sonnet-5-5"

spend_limits:
  defaults:
    user_daily_usd: 10
    team_monthly_usd: 1000
  on_exceeded: "deny"

Put OIDC secrets, database URL, upstream keys and telemetry headers in environment variables or secret management, not plaintext gateway.yaml.

#Generic LLM gateway protocol

Code still sends Anthropic Messages through Passion8/custom gateways. Set the root in ANTHROPIC_BASE_URL and provide compatible paths.

ItemRequirement
Base URLHost root such as https://passion8.cc, usually without /v1
MessagesPOST /v1/messages required
Token countOptional; absence causes estimates or reduced functionality
StreamingSSE for stream:true without proxy buffering
HeadersPreserve version/beta headers rather than forcing old values
ErrorsPreserve JSON status, type and message for retries/fallback

Use extensible validation rather than a minimal fixed schema. Preserve these fields and explicitly forward or reject new beta fields:

FieldMeaning
model, max_tokens, messages, systemCore Messages fields
streamSSE control
tools, tool_choiceBuilt-in/MCP/deferred tools
thinkingReasoning/effort
metadata, stop_sequencesBusiness metadata and termination
temperature, top_p, top_kSampling
context_management, output_configContext and output control
cache_controlNested system/message/tool blocks

#System arrays and attribution

Preserve system array block order, types and nested fields, including the attribution block's position.

AvoidRisk
Flattening arraysBreaks attribution, cache keys and controls
Prepending custom blocksAttribution may enter the actual prompt and prefix
Sorting/deduplicating/rewritingCache misses and policy/capability errors
Dropping unknown fieldsNew beta features fail

Prefer managed/provider policy. If rewriting is unavoidable, preserve attribution ordering and audit the transformation.

#Model discovery

Enable gateway models in the picker with:

export CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1

Discovery request:

GET /v1/models?limit=1000

Operational constraints:

ItemDetails
TimeoutAbout three seconds; slow responses, redirects or OIDC cause silent failure
CacheUsually ~/.claude/cache/gateway-models.json
Manual defaultsANTHROPIC_MODEL and DEFAULT_OPUS/SONNET/HAIKU/FABLE_MODEL variables
CapabilitiesListing does not establish search, 1M context, thinking or one-hour caching

Return stable, lightweight, noninteractive model JSON; never redirect discovery to an IdP or login page.

#One-hour prompt caching

A client variable alone cannot guarantee one-hour caching. Verify:

ConditionCheck
ClientEnable one-hour TTL without FORCE_PROMPT_CACHING_5M or DISABLE_PROMPT_CACHING
UpstreamActual routed provider/model supports it
HeadersNo version/beta stripping or downgrade
Stable bodyPreserve arrays, messages, tools, cache controls and model parameters
UsagePreserve creation/read token fields

Document and observe five-minute fallback if one hour is unsupported. Never fabricate hits through body rewriting.

#Spend limits

Consider budgets at gateway and provider levels:

LayerRecommendation
GatewayDaily/monthly limits and concurrency by user/team/project/model/upstream
ProviderHard caps against bugs or credential exposure
ResponseStable JSON 402/403/429 with limit, window and reset
ObservabilitySpend, denials, matched rules and fallback

Do not silently substitute a cheaper, different model after a limit unless policy and UI explicitly explain it. Silent fallback complicates debugging, cost attribution and auditing.

#Server-managed settings limitations

Official settings depend on Anthropic identities and policy services. Do not assume they govern third-party providers or deliver custom gateway URLs/tokens.

ScenarioRecommendation
Standard URLMDM, endpoint/system settings or templates
Token deliveryapiKeyHelper, vault or short-lived credentials; no repository secrets
Models/budgetsGateway policies and spend limits
Apps gatewayIts OIDC, Postgres, managed policy and telemetry

#Passion8/custom gateway checklist

CheckDetails
Base URLRoot such as https://passion8.cc, not OpenAI /v1
AuthUsually ANTHROPIC_AUTH_TOKEN; avoid stale API-key/login conflicts
Private accessLogin, callback and admin/health reachable on controlled network
ProxyUnbuffered SSE, persistent connections and suitable idle timeout
ForwardingPreserve headers, arrays, tools, caching, usage and errors
AuditRequest ID, user, model, tokens and policy; no full payload by default

#curl self-checks

claude-sonnet-5-5 is a current official example ID. Verify Passion8 access first or substitute an enabled console model.

Set environment variables first:

export ANTHROPIC_BASE_URL="https://passion8.cc"
export ANTHROPIC_AUTH_TOKEN="sk-..."
export ANTHROPIC_MODEL="claude-sonnet-5-5"

Messages JSON:

curl -sS "$ANTHROPIC_BASE_URL/v1/messages" \
  -H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "'"$ANTHROPIC_MODEL"'",
    "max_tokens": 64,
    "messages": [
      { "role": "user", "content": "Reply only ok" }
    ]
  }'

SSE stream:

curl -N "$ANTHROPIC_BASE_URL/v1/messages" \
  -H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "'"$ANTHROPIC_MODEL"'",
    "max_tokens": 64,
    "stream": true,
    "messages": [
      { "role": "user", "content": "stream ok" }
    ]
  }'

Test discovery within the client-like three-second budget:

curl -sS --max-time 3 "$ANTHROPIC_BASE_URL/v1/models?limit=1000" \
  -H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN" \
  -H "anthropic-version: 2023-06-01"

Optional token count:

curl -sS "$ANTHROPIC_BASE_URL/v1/messages/count_tokens" \
  -H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "'"$ANTHROPIC_MODEL"'",
    "messages": [
      { "role": "user", "content": "count tokens" }
    ]
  }'

Expected checks:

ItemPassing condition
MessagesAnthropic JSON, not HTML/login/proxy pages
SSEcurl -N receives event/data incrementally without buffering
ModelsJSON or explicit JSON 401/403 within three seconds
Count tokensReal count when supported, explicit error otherwise
Spend capStable JSON error with attributable telemetry

#Official references

Support

Need help?

For setup, billing, or model issues, email us. Check the status page for uptime.

WeChat / QQ support is available at the bottom right.