Claude Code

Gateway protocol deployment checklist

Protocols, forwarding, discovery, SSO, Postgres, spend limits, error semantics and cache acceptance for LLM and apps gateways.

Gateway failures often involve headers, bodies, streams, discovery, usage and errors rather than just the URL. This page turns official protocol and rollout guidance into deployment checks.

Do not freeze a field allowlist around today's requests. New beta headers, schemas and context fields can otherwise silently degrade or fail with 400.

#Three gateway forms

FormClient configurationImplementation
Anthropic MessagesANTHROPIC_BASE_URL/v1/messages SSE; optional count_tokens and models
Bedrock formatCLAUDE_CODE_USE_BEDROCK=1 plus ANTHROPIC_BEDROCK_BASE_URLInvokeModel and streaming paths with dialect translation
Vertex/Agent PlatformCLAUDE_CODE_USE_VERTEX=1 plus ANTHROPIC_VERTEX_BASE_URLrawPredict, streamRawPredict and count-token paths

Passion8 and most Code gateways use Anthropic Messages. Do not confuse this with OpenAI chat/completions.

#Minimum protocol requirements

ItemRequirementFailure symptom
StreamingForward Messages as SSEClient appears stuck
anthropic-versionPreserve unchangedVersion/schema mismatch
anthropic-betaPreserve without a fixed filterSearch, context management and extended-context features break
system arrayPreserve block order and structureAttribution, prompts and cache keys break
toolsPreserve schemas and deferred-search fieldsTool failures and repeated cache misses
thinking/output_configSupport or explicitly translateReasoning, effort and structured-output errors
Error envelopePreserve status, type, message and request IDIncorrect retry/fallback
UsagePreserve input/output/cache tokensMisleading usage and cache diagnostics

#Credentials and Base URL

Typical gateway configuration:

export ANTHROPIC_BASE_URL="https://llm-gateway.example.com"
export ANTHROPIC_AUTH_TOKEN="sk-gateway-token"

If the gateway requires x-api-key:

export ANTHROPIC_API_KEY="sk-gateway-key"

Credential choices:

CaseRecommendation
Fixed personal tokenUser settings environment
Rotating tokenapiKeyHelper or enterprise helper
Tenant routingCustom headers with gateway-side validation
Repository configurationNonsecret defaults only

#Attribution blocks and caching

Code inserts an attribution block into the system prompt. The official endpoint strips it when correctly positioned to protect first-party caching. Third-party gateways should preserve the array or account explicitly for cache consequences.

Gateway behaviorCache consequence
Preserve array with attribution firstClosest to the official path
Prepend company policyAttribution may enter the prompt and cache key
Flatten to stringBreaks stripping and increases misses
Rewrite/reorder tool schemasUnstable prefix

Prefer managed settings, CLAUDE.md, permissions or provider policy over dynamically modifying every system prefix.

#Model discovery

Optional /v1/models enables the model picker:

export CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1

Deployment checks:

CheckRequirement
PathGET /v1/models?limit=1000 returns promptly
IDsDirectly usable by Claude Code
CapabilitiesPicker presence does not prove effort, fast mode, search or 1M support
CacheClear local discovery cache during diagnostics
ErrorsSlow responses, redirects or login HTML may be silently ignored

#Apps-gateway configuration areas

The official claude gateway --config gateway.yaml form handles organizational SSO, policies and multiple cloud upstreams.

AreaPurposeRisk
serverListener, public URL, TLSDevices and browser callbacks must reach it
auth.oidcIdP, client, claims, groupsBad callbacks/scopes/claims cause login loops
sessionBearer signing and TTLShort TTL without refresh causes frequent login
storePostgres grants and limitsReplicas need shared database and planned migrations
upstreamsAnthropic, Bedrock, Agent Platform, FoundryID mismatch causes 404 or bad fallback
managedSettings, model policy, permissionsDo not assume official settings cover third-party paths
telemetryOTLP metrics/traces/logsAvoid raw prompts and tool payloads by default
adminSpend limits and Admin APIRotate separate keys by purpose/environment

#Spend limits

Spend limits are gateway circuit breakers, not authoritative bills. They estimate daily/weekly/monthly developer spend from streamed usage and pricing.

DesignMeaning
Scopeuser、rbac_group、organization
AmountCents string; null unlimited, zero blocks
Perioddaily、weekly、monthly
Effective capUser override, then group, then organization
Over limit429, billing_error, x-should-retry: false
count_tokensUsually unbilled and not blocked by caps
Provider billActual cloud/Anthropic bill remains authoritative

Choose fail-open or fail-closed on database outages according to availability versus strict-budget requirements.

#Error semantics

ErrorPreserve/returnClient behavior
401/403Clear authentication messageCheck credentials/login/upstream access
404 modelModel unavailable on this upstreamFallback or model change
429Rate/spend limit with retry semanticsWait or stop retrying
5xx/timeoutTransient upstream failureRetry or switch upstream
400 unsupported fieldIdentify thinking/context_management/etc.Disable beta temporarily or change provider
Login HTMLNever return it on API pathsOtherwise malformed-response failure

#Rollout order

  1. Verify URL/token with one user's shell configuration.
  2. Fix model/effort, ask follow-ups without MCP and inspect caching.
  3. Add discovery and verify picker/fallback.
  4. Add MCP, tool search and plugins; inspect schema stability.
  5. Deliver user or endpoint-managed settings.
  6. Add OTel with session/agent/model/team dimensions.
  7. Configure spend limits or budget alerts.
  8. Exercise failover and error semantics.
  9. Expand to desktop, IDE, CI and SDK last.

#Cache acceptance

CheckPassing condition
Stable model/effort follow-upsCache reads appear from subsequent turns
Five-minute pauseVerify behavior at the provider's short-TTL boundary
One-hour eligible planLonger intervals can reuse subscription cache
Many MCP toolsSearch/deferred loading avoids all-schema repetition
SubagentsIndependent caches correctly attributed
Gateway usageCreation/read tokens forwarded unchanged

#Official references

Support

Need help?

For setup, billing, or model issues, email us. Check the status page for uptime.

WeChat / QQ support is available at the bottom right.