OpenAI · Python
The OpenAI Python SDK accepts a base_url argument on the client constructor.
Point it at the gateway and every chat completion routes to a model
from your connected providers without changing the call sites in your application code. This
recipe ships a 30-line script that you can drop into a fresh virtualenv,
install one package, and run against your own org.
Drop-in switch
Two lines change in your existing client construction.
client = OpenAI(
base_url="https://gateway.iq-routing.com/v1",
api_key="gw_live_xxxxxxxx",
)
See Authentication for how to generate a gw_live_...
key and where to store it. The base_url swap routes the request through the gateway. Everything below
the client (calls to client.chat.completions.create and
client.models.list) hits the gateway, which forwards to the upstream provider
the gateway picks. The gateway serves the chat, models, and Responses paths;
it does not implement /v1/embeddings, so leave embedding calls pointed at
your provider.
Verify it routes
import json
import os
from openai import OpenAI
client = OpenAI(
base_url="https://gateway.iq-routing.com/v1",
api_key=os.environ["IQ_API_KEY"],
)
response = client.chat.completions.with_raw_response.create(
model="auto",
messages=[{"role": "user", "content": "Summarise: routing-aware gateway."}],
max_tokens=64,
)
print("X-Request-Id:", response.headers.get("x-request-id"))
telemetry = json.loads(response.headers.get("x-iq-routing", "{}"))
print("Model picked:", telemetry.get("chosen_model"))
print("Provider:", telemetry.get("chosen_provider"))
print("Cache hit:", telemetry.get("cache_hit"))
print(response.parse().choices[0].message.content)
Set IQ_API_KEY in your environment, run the script. The gateway emits
two response headers per request: X-Request-Id (the correlation handle)
and x-iq-routing (a JSON-encoded telemetry blob with the chosen model,
provider, complexity tier, cache-hit flag, latencies, and cost). Paste
the X-Request-Id value into /requests/<uuid> to see the full
routing decision (fallback hops, the provider's raw
response).
The model="auto" argument picks a model from your connected providers, balancing quality and cost. Pinning a specific
model (gpt-4o-mini, claude-haiku-4-5) keeps the request with that
provider; name an exact-pin id (gpt-6-astra, gpt-6.1-sol,
gpt-6-sol, claude-opus-5-5, claude-fable-5-1) for an exact model. See
Models for GPT-6.1 Sol request settings and pricing.
Using capability aliases
The cap:<name> syntax in the model field tells the gateway to pick a
concrete model at routing time based on the capability's stable
intent rather than a pinned model id. Your org's per-capability
overrides and current provider health can shift which concrete model a
capability resolves to on any given call; the capability's intent
stays stable (reason-heavy always favors a stronger reasoning model).
Five default capabilities resolve to a concrete model this way:
reason-heavy, tool-call-strict, long-context-128k, vision, and
json-mode. A sixth, cheap-fast, is accepted too but currently
routes like auto -- see capability aliases.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://gateway.iq-routing.com/v1",
api_key=os.environ["IQ_API_KEY"],
)
completion = client.chat.completions.create(
model="cap:reason-heavy",
messages=[{"role": "user", "content": "Plan the migration in three phases."}],
)
The resolved provider plus model surface in the x-iq-routing response
header (a JSON-encoded payload with chosen_provider and chosen_model
fields, among others), so the SDK consumer can inspect which concrete
model handled the request via with_raw_response. Decode the header with
json.loads(response.headers["x-iq-routing"]). See the capability
aliases docs for the full list of default
mappings and how to override them. You can override the default
capability mapping for your org in the dashboard at
/settings/capabilities.
Common gotchas
A missing or invalid API key returns 401 before the request reaches
routing -- no provider call is made and nothing is billed. See the
API reference for the full set of status codes
and error body shapes.
The OpenAI Python SDK's default request timeout is 600 seconds,
which matches the gateway's own per-provider attempt timeout -- a single
hung provider call will not outlast the SDK's default. The gateway's
outer ceiling across a full fallback chain is 1800 seconds (30 minutes);
past that it returns 504 with a request_id in the body. Most requests
finish in a few seconds. If your product needs to fail faster than either of
those worst cases, set an explicit timeout on the client:
client = OpenAI(
base_url="https://gateway.iq-routing.com/v1",
api_key=os.environ["IQ_API_KEY"],
timeout=60.0,
)
When every upstream provider in the fallback chain is exhausted before
that outer ceiling is hit, the gateway returns 503 with an
overloaded_error type in the response body, and no Retry-After
header -- back off on your own schedule for that case.
The SDK's default retry behaviour catches 429 and retries with
exponential backoff up to two times. The gateway's 429 already
encodes a Retry-After header set to the exact cooldown; SDK-side
retries on top of that just burn calls inside the same cooldown
window. Disable SDK retries when you call through the gateway:
client = OpenAI(
base_url="https://gateway.iq-routing.com/v1",
api_key=os.environ["IQ_API_KEY"],
timeout=60.0,
max_retries=0,
)
Streaming works as you would expect against /v1/chat/completions. The
gateway forwards the SSE stream chunk-by-chunk; the X-Request-Id header
arrives on the first event so the correlation handle is available before the
last token.