Grok

Grok API and SDK

Call Grok with a Passion8 API key and /v1 URL. Includes OpenAI Python SDK examples for Chat Completions, Responses and streaming, plus reasoning and search-tool limits.

Use the OpenAI SDK to call Passion8 Grok from your own application. Prepare a Passion8 key and an enabled console model. This guide starts with a minimal Chat Completions call, then distinguishes Responses and official hosted tools.

#Configure URL, key and model

SettingValue
Base URLhttps://passion8.cc/v1
API KeyA key created in the Passion8 console
Example modelgrok-4.7, after confirming account access

macOS / Linux:

export PASSION8_API_KEY="replace-with-your-Passion8-key"
export PASSION8_GROK_MODEL="grok-4.7"

Windows PowerShell:

$env:PASSION8_API_KEY="replace-with-your-Passion8-key"
$env:PASSION8_GROK_MODEL="grok-4.7"

If your console provides another model, use its actual ID. The following code reads the key from the environment and does not load Grok Build or Codex configuration.

#First request: Chat Completions

Install in your project virtual environment:

python -m pip install --upgrade openai

Save as grok_smoke.py:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["PASSION8_API_KEY"],
    base_url="https://passion8.cc/v1",
)
response = client.chat.completions.create(
    model=os.environ["PASSION8_GROK_MODEL"],
    messages=[{"role": "user", "content": "Reply only OK."}],
)
print(response.choices[0].message.content)

Run python grok_smoke.py and check model and request time in the console. Asking a model to identify itself does not verify routing.

#Stream output

After the client initialization, use:

stream = client.chat.completions.create(
    model=os.environ["PASSION8_GROK_MODEL"],
    messages=[{"role": "user", "content": "Explain the purpose of unit tests."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="", flush=True)
print()

Streaming returns incremental chunks. Keep reading instead of parsing the entire response as a single JSON object. Verify billing in console records.

#Responses: choose according to the client

xAI currently recommends Responses for new integrations and labels Chat Completions as legacy. For existing coding extensions, configure the protocol the client actually sends.

If your account's channel supports Passion8 Grok Responses, replace the request with:

response = client.responses.create(
    model=os.environ["PASSION8_GROK_MODEL"],
    input="Reply only OK.",
)
print(response.output_text)

Responses uses input and output_text; Chat Completions uses messages and choices. Change response parsing along with the API method. Confirm support for upstream previous_response_id, background execution and other features separately; a text response does not establish their availability.

#Reasoning, tools and context

Official Grok 4.7 has a 500,000-token context window and supports low, medium, high and xhigh, with default high. Reasoning parameters differ by protocol. Add them using that protocol's documentation after a minimal call succeeds; do not send the old none value.

Function calling returns a tool request for your application to execute and return. Official Web Search, X Search and Code Execution are hosted tools. Do not copy full examples containing hosted tools before confirming channel support. Context size is also different from maximum output tokens.

#Troubleshooting

ErrorFirst checks
401 / 403Passion8 key, token status, model access and balance
404/v1 base URL, request protocol and model ID
400 invalid parameterMixed Chat / Responses fields or unsupported tools
429Error details, concurrency and rate-limit guidance
Context overflowClient budget and actual channel limit; do not resend the same oversized request repeatedly

For editor use, see Grok in VS Code. For the official client, see Grok Build.

#API Setup Questions

#Do Cline and Codex use the same Grok API endpoint?

Both can use https://passion8.cc/v1 as the base URL, but Cline OpenAI Compatible uses Chat Completions, while this guide's Codex configuration uses Responses. A shared base URL does not mean identical request bodies or support for both protocols on your account route. See Cline multi-model setup and the Grok setup guide for the corresponding editor instructions.

#Does a successful Grok text response prove that X or web search works?

No. Text generation, function calling and xAI-hosted X Search / Web Search are separate capabilities. Use hosted tools only when the current route explicitly supports them. A text answer or successful Cline local-tool execution does not confirm hosted search availability.

#Verified sources

Checked: 2026-10-11.

Passion8 model and channel support depend on console access and actual error responses.

Support

Need help?

For setup, billing, or model issues, email us. Check the status page for uptime.

WeChat / QQ support is available at the bottom right.