# Claude Code Gateway protocol deployment checklist

> Claude Code: Protocols, forwarding, discovery, SSO, Postgres, spend limits, error semantics and cache acceptance for LLM and apps gateways.

URL: https://docs.passion8.cc/en/docs/claude-code/gateway-protocol
Language: en
Publisher: Passion8

Gateway failures often involve headers, bodies, streams, discovery, usage and errors rather than just the URL. This page turns official protocol and rollout guidance into deployment checks.




Do not freeze a field allowlist around today's requests. New beta headers, schemas and context fields can otherwise silently degrade or fail with 400.




## Three gateway forms

| Form | Client configuration | Implementation |
| --- | --- | --- |
| Anthropic Messages | ANTHROPIC_BASE_URL | /v1/messages SSE; optional count_tokens and models |
| Bedrock format | CLAUDE_CODE_USE_BEDROCK=1 plus ANTHROPIC_BEDROCK_BASE_URL | InvokeModel and streaming paths with dialect translation |
| Vertex/Agent Platform | CLAUDE_CODE_USE_VERTEX=1 plus ANTHROPIC_VERTEX_BASE_URL | rawPredict, streamRawPredict and count-token paths |

Passion8 and most Code gateways use Anthropic Messages. Do not confuse this with OpenAI chat/completions.

## Minimum protocol requirements

| Item | Requirement | Failure symptom |
| --- | --- | --- |
| Streaming | Forward Messages as SSE | Client appears stuck |
| anthropic-version | Preserve unchanged | Version/schema mismatch |
| anthropic-beta | Preserve without a fixed filter | Search, context management and extended-context features break |
| system array | Preserve block order and structure | Attribution, prompts and cache keys break |
| tools | Preserve schemas and deferred-search fields | Tool failures and repeated cache misses |
| thinking/output_config | Support or explicitly translate | Reasoning, effort and structured-output errors |
| Error envelope | Preserve status, type, message and request ID | Incorrect retry/fallback |
| Usage | Preserve input/output/cache tokens | Misleading usage and cache diagnostics |

## Credentials and Base URL

Typical gateway configuration:

```bash
export ANTHROPIC_BASE_URL="https://llm-gateway.example.com"
export ANTHROPIC_AUTH_TOKEN="sk-gateway-token"
```

If the gateway requires x-api-key:

```bash
export ANTHROPIC_API_KEY="sk-gateway-key"
```

Credential choices:

| Case | Recommendation |
| --- | --- |
| Fixed personal token | User settings environment |
| Rotating token | apiKeyHelper or enterprise helper |
| Tenant routing | Custom headers with gateway-side validation |
| Repository configuration | Nonsecret defaults only |

## Attribution blocks and caching

Code inserts an attribution block into the system prompt. The official endpoint strips it when correctly positioned to protect first-party caching. Third-party gateways should preserve the array or account explicitly for cache consequences.

| Gateway behavior | Cache consequence |
| --- | --- |
| Preserve array with attribution first | Closest to the official path |
| Prepend company policy | Attribution may enter the prompt and cache key |
| Flatten to string | Breaks stripping and increases misses |
| Rewrite/reorder tool schemas | Unstable prefix |

Prefer managed settings, CLAUDE.md, permissions or provider policy over dynamically modifying every system prefix.

## Model discovery

Optional /v1/models enables the model picker:

```bash
export CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1
```

Deployment checks:

| Check | Requirement |
| --- | --- |
| Path | GET /v1/models?limit=1000 returns promptly |
| IDs | Directly usable by Claude Code |
| Capabilities | Picker presence does not prove effort, fast mode, search or 1M support |
| Cache | Clear local discovery cache during diagnostics |
| Errors | Slow responses, redirects or login HTML may be silently ignored |

## Apps-gateway configuration areas

The official claude gateway --config gateway.yaml form handles organizational SSO, policies and multiple cloud upstreams.

| Area | Purpose | Risk |
| --- | --- | --- |
| server | Listener, public URL, TLS | Devices and browser callbacks must reach it |
| auth.oidc | IdP, client, claims, groups | Bad callbacks/scopes/claims cause login loops |
| session | Bearer signing and TTL | Short TTL without refresh causes frequent login |
| store | Postgres grants and limits | Replicas need shared database and planned migrations |
| upstreams | Anthropic, Bedrock, Agent Platform, Foundry | ID mismatch causes 404 or bad fallback |
| managed | Settings, model policy, permissions | Do not assume official settings cover third-party paths |
| telemetry | OTLP metrics/traces/logs | Avoid raw prompts and tool payloads by default |
| admin | Spend limits and Admin API | Rotate separate keys by purpose/environment |

## Spend limits

Spend limits are gateway circuit breakers, not authoritative bills. They estimate daily/weekly/monthly developer spend from streamed usage and pricing.

| Design | Meaning |
| --- | --- |
| Scope | `user`、`rbac_group`、`organization` |
| Amount | Cents string; null unlimited, zero blocks |
| Period | `daily`、`weekly`、`monthly` |
| Effective cap | User override, then group, then organization |
| Over limit | 429, billing_error, x-should-retry: false |
| count_tokens | Usually unbilled and not blocked by caps |
| Provider bill | Actual cloud/Anthropic bill remains authoritative |

Choose fail-open or fail-closed on database outages according to availability versus strict-budget requirements.

## Error semantics

| Error | Preserve/return | Client behavior |
| --- | --- | --- |
| 401/403 | Clear authentication message | Check credentials/login/upstream access |
| 404 model | Model unavailable on this upstream | Fallback or model change |
| 429 | Rate/spend limit with retry semantics | Wait or stop retrying |
| 5xx/timeout | Transient upstream failure | Retry or switch upstream |
| 400 unsupported field | Identify thinking/context_management/etc. | Disable beta temporarily or change provider |
| Login HTML | Never return it on API paths | Otherwise malformed-response failure |

## Rollout order

1. Verify URL/token with one user's shell configuration.
2. Fix model/effort, ask follow-ups without MCP and inspect caching.
3. Add discovery and verify picker/fallback.
4. Add MCP, tool search and plugins; inspect schema stability.
5. Deliver user or endpoint-managed settings.
6. Add OTel with session/agent/model/team dimensions.
7. Configure spend limits or budget alerts.
8. Exercise failover and error semantics.
9. Expand to desktop, IDE, CI and SDK last.

## Cache acceptance

| Check | Passing condition |
| --- | --- |
| Stable model/effort follow-ups | Cache reads appear from subsequent turns |
| Five-minute pause | Verify behavior at the provider's short-TTL boundary |
| One-hour eligible plan | Longer intervals can reuse subscription cache |
| Many MCP tools | Search/deferred loading avoids all-schema repetition |
| Subagents | Independent caches correctly attributed |
| Gateway usage | Creation/read tokens forwarded unchanged |

## Official references

- [Gateway protocol reference](https://code.claude.com/docs/en/llm-gateway-protocol.md)
- [Connect Claude Code to an LLM gateway](https://code.claude.com/docs/en/llm-gateway-connect.md)
- [Roll out an LLM gateway](https://code.claude.com/docs/en/llm-gateway-rollout.md)
- [Claude apps gateway configuration](https://code.claude.com/docs/en/claude-apps-gateway-config.md)
- [Claude apps gateway spend limits](https://code.claude.com/docs/en/claude-apps-gateway-spend-limits.md)
