# Grok API Tutorial: OpenAI Python SDK, Responses and Streaming

> Call Grok with a Passion8 API key and /v1 URL. Includes OpenAI Python SDK examples for Chat Completions, Responses and streaming, plus reasoning and search-tool limits.

URL: https://docs.passion8.cc/en/docs/grok/api
Language: en
Publisher: Passion8
Last updated: 2026-10-11

Use the OpenAI SDK to call Passion8 Grok from your own application. Prepare a Passion8 key and an enabled console model. This guide starts with a minimal Chat Completions call, then distinguishes Responses and official hosted tools.

## Configure URL, key and model

| Setting | Value |
| --- | --- |
| Base URL | `https://passion8.cc/v1` |
| API Key | A key created in the Passion8 console |
| Example model | `grok-4.7`, after confirming account access |

macOS / Linux:

```bash
export PASSION8_API_KEY="replace-with-your-Passion8-key"
export PASSION8_GROK_MODEL="grok-4.7"
```

Windows PowerShell:

```powershell
$env:PASSION8_API_KEY="replace-with-your-Passion8-key"
$env:PASSION8_GROK_MODEL="grok-4.7"
```

If your console provides another model, use its actual ID. The following code reads the key from the environment and does not load Grok Build or Codex configuration.

## First request: Chat Completions

Install in your project virtual environment:

```bash
python -m pip install --upgrade openai
```

Save as `grok_smoke.py`:

```python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["PASSION8_API_KEY"],
    base_url="https://passion8.cc/v1",
)
response = client.chat.completions.create(
    model=os.environ["PASSION8_GROK_MODEL"],
    messages=[{"role": "user", "content": "Reply only OK."}],
)
print(response.choices[0].message.content)
```

Run `python grok_smoke.py` and check model and request time in the [console](https://passion8.cc). Asking a model to identify itself does not verify routing.

## Stream output

After the client initialization, use:

```python
stream = client.chat.completions.create(
    model=os.environ["PASSION8_GROK_MODEL"],
    messages=[{"role": "user", "content": "Explain the purpose of unit tests."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="", flush=True)
print()
```

Streaming returns incremental chunks. Keep reading instead of parsing the entire response as a single JSON object. Verify billing in console records.

## Responses: choose according to the client

xAI currently recommends Responses for new integrations and labels Chat Completions as legacy. For existing coding extensions, configure the protocol the client actually sends.

If your account's channel supports Passion8 Grok Responses, replace the request with:

```python
response = client.responses.create(
    model=os.environ["PASSION8_GROK_MODEL"],
    input="Reply only OK.",
)
print(response.output_text)
```

Responses uses `input` and `output_text`; Chat Completions uses `messages` and `choices`. Change response parsing along with the API method. Confirm support for upstream `previous_response_id`, background execution and other features separately; a text response does not establish their availability.

## Reasoning, tools and context

Official Grok 4.7 has a **500,000-token** context window and supports `low`, `medium`, `high` and `xhigh`, with default `high`. Reasoning parameters differ by protocol. Add them using that protocol's documentation after a minimal call succeeds; do not send the old `none` value.

Function calling returns a tool request for your application to execute and return. Official Web Search, X Search and Code Execution are hosted tools. Do not copy full examples containing hosted tools before confirming channel support. Context size is also different from maximum output tokens.

## Troubleshooting

| Error | First checks |
| --- | --- |
| 401 / 403 | Passion8 key, token status, model access and balance |
| 404 | `/v1` base URL, request protocol and model ID |
| 400 invalid parameter | Mixed Chat / Responses fields or unsupported tools |
| 429 | Error details, concurrency and rate-limit guidance |
| Context overflow | Client budget and actual channel limit; do not resend the same oversized request repeatedly |

For editor use, see [Grok in VS Code](https://docs.passion8.cc/en/docs/grok/vscode). For the official client, see [Grok Build](https://docs.passion8.cc/en/docs/grok).

## API Setup Questions

### Do Cline and Codex use the same Grok API endpoint?

Both can use `https://passion8.cc/v1` as the base URL, but Cline OpenAI Compatible uses Chat Completions, while this guide's Codex configuration uses Responses. A shared base URL does not mean identical request bodies or support for both protocols on your account route. See [Cline multi-model setup](https://docs.passion8.cc/en/docs/cline) and the [Grok setup guide](https://docs.passion8.cc/en/docs/grok) for the corresponding editor instructions.

### Does a successful Grok text response prove that X or web search works?

No. Text generation, function calling and xAI-hosted X Search / Web Search are separate capabilities. Use hosted tools only when the current route explicitly supports them. A text answer or successful Cline local-tool execution does not confirm hosted search availability.

## Verified sources

Checked: 2026-10-11.

- [Official Grok text generation](https://docs.x.ai/developers/model-capabilities/text/generate-text)
- [Chat Completions legacy](https://docs.x.ai/developers/model-capabilities/legacy/chat-completions)
- [Grok 4.7 capabilities and reasoning](https://docs.x.ai/developers/models/grok-4.7)

Passion8 model and channel support depend on console access and actual error responses.
