Skip to content
IQ Routing

OpenAI · Python

The OpenAI Python SDK accepts a base_url argument on the client constructor. Point it at the gateway and every chat completion routes to a model from your connected providers without changing the call sites in your application code. This recipe ships a 30-line script that you can drop into a fresh virtualenv, install one package, and run against your own org.

Drop-in switch

Two lines change in your existing client construction.

client = OpenAI(
    base_url="https://gateway.iq-routing.com/v1",
    api_key="gw_live_xxxxxxxx",
)

See Authentication for how to generate a gw_live_... key and where to store it. The base_url swap routes the request through the gateway. Everything below the client (calls to client.chat.completions.create and client.models.list) hits the gateway, which forwards to the upstream provider the gateway picks. The gateway serves the chat, models, and Responses paths; it does not implement /v1/embeddings, so leave embedding calls pointed at your provider.

Verify it routes

import json
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://gateway.iq-routing.com/v1",
    api_key=os.environ["IQ_API_KEY"],
)

response = client.chat.completions.with_raw_response.create(
    model="auto",
    messages=[{"role": "user", "content": "Summarise: routing-aware gateway."}],
    max_tokens=64,
)

print("X-Request-Id:", response.headers.get("x-request-id"))
telemetry = json.loads(response.headers.get("x-iq-routing", "{}"))
print("Model picked:", telemetry.get("chosen_model"))
print("Provider:", telemetry.get("chosen_provider"))
print("Cache hit:", telemetry.get("cache_hit"))
print(response.parse().choices[0].message.content)

Set IQ_API_KEY in your environment, run the script. The gateway emits two response headers per request: X-Request-Id (the correlation handle) and x-iq-routing (a JSON-encoded telemetry blob with the chosen model, provider, complexity tier, cache-hit flag, latencies, and cost). Paste the X-Request-Id value into /requests/<uuid> to see the full routing decision (fallback hops, the provider's raw response).

The model="auto" argument picks a model from your connected providers, balancing quality and cost. Pinning a specific model (gpt-4o-mini, claude-haiku-4-5) keeps the request with that provider; name an exact-pin id (gpt-6-astra, gpt-6.1-sol, gpt-6-sol, claude-opus-5-5, claude-fable-5-1) for an exact model. See Models for GPT-6.1 Sol request settings and pricing.

Using capability aliases

The cap:<name> syntax in the model field tells the gateway to pick a concrete model at routing time based on the capability's stable intent rather than a pinned model id. Your org's per-capability overrides and current provider health can shift which concrete model a capability resolves to on any given call; the capability's intent stays stable (reason-heavy always favors a stronger reasoning model). Five default capabilities resolve to a concrete model this way: reason-heavy, tool-call-strict, long-context-128k, vision, and json-mode. A sixth, cheap-fast, is accepted too but currently routes like auto -- see capability aliases.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://gateway.iq-routing.com/v1",
    api_key=os.environ["IQ_API_KEY"],
)

completion = client.chat.completions.create(
    model="cap:reason-heavy",
    messages=[{"role": "user", "content": "Plan the migration in three phases."}],
)

The resolved provider plus model surface in the x-iq-routing response header (a JSON-encoded payload with chosen_provider and chosen_model fields, among others), so the SDK consumer can inspect which concrete model handled the request via with_raw_response. Decode the header with json.loads(response.headers["x-iq-routing"]). See the capability aliases docs for the full list of default mappings and how to override them. You can override the default capability mapping for your org in the dashboard at /settings/capabilities.

Common gotchas

A missing or invalid API key returns 401 before the request reaches routing -- no provider call is made and nothing is billed. See the API reference for the full set of status codes and error body shapes.

The OpenAI Python SDK's default request timeout is 600 seconds, which matches the gateway's own per-provider attempt timeout -- a single hung provider call will not outlast the SDK's default. The gateway's outer ceiling across a full fallback chain is 1800 seconds (30 minutes); past that it returns 504 with a request_id in the body. Most requests finish in a few seconds. If your product needs to fail faster than either of those worst cases, set an explicit timeout on the client:

client = OpenAI(
    base_url="https://gateway.iq-routing.com/v1",
    api_key=os.environ["IQ_API_KEY"],
    timeout=60.0,
)

When every upstream provider in the fallback chain is exhausted before that outer ceiling is hit, the gateway returns 503 with an overloaded_error type in the response body, and no Retry-After header -- back off on your own schedule for that case.

The SDK's default retry behaviour catches 429 and retries with exponential backoff up to two times. The gateway's 429 already encodes a Retry-After header set to the exact cooldown; SDK-side retries on top of that just burn calls inside the same cooldown window. Disable SDK retries when you call through the gateway:

client = OpenAI(
    base_url="https://gateway.iq-routing.com/v1",
    api_key=os.environ["IQ_API_KEY"],
    timeout=60.0,
    max_retries=0,
)

Streaming works as you would expect against /v1/chat/completions. The gateway forwards the SSE stream chunk-by-chunk; the X-Request-Id header arrives on the first event so the correlation handle is available before the last token.

Try this in the playground

Run this snippet against your gateway key →