Anthropic · Python
The Anthropic Python SDK accepts a base_url argument on the client. Point
it at the gateway and the standard client.messages.create calls route
through. This recipe pins claude-haiku-4-5-20251001 directly, which
keeps every call in the example with Anthropic; name claude-opus-5-5
or claude-fable-5-1 instead for an exact model. See the Using
capability aliases section below, or pass model="auto", to let the
gateway pick the model per request instead.
Drop-in switch
client = anthropic.Anthropic(
base_url="https://gateway.iq-routing.com",
api_key="gw_live_xxxxxxxx",
)
The gateway exposes the Anthropic surface at /v1/messages (no /v1
prefix in the SDK's internal routing; the SDK appends /v1/messages
itself). Pass the gateway origin without /v1.
Verify it routes
import os
import anthropic
client = anthropic.Anthropic(
base_url="https://gateway.iq-routing.com",
api_key=os.environ["IQ_API_KEY"],
timeout=60.0,
max_retries=0,
)
response = client.messages.with_raw_response.create(
model="claude-haiku-4-5-20251001",
max_tokens=64,
messages=[{"role": "user", "content": "Summarise: routing-aware gateway."}],
)
print("X-Request-Id:", response.headers.get("x-request-id"))
parsed = response.parse()
print("Stop reason:", parsed.stop_reason)
print(parsed.content[0].text)
Drop the script into a fresh venv, pip install anthropic, set
IQ_API_KEY in your environment, run. The X-Request-Id header is the
gateway's correlation handle.
The model pin (claude-haiku-4-5-20251001) keeps every call with
Anthropic; name claude-opus-5-5 or claude-fable-5-1 instead for an
exact model. If you pass model="auto" the gateway picks a model
from your connected providers instead, balancing quality and cost. The
auto pick is cross-provider, so a turn can land on an OpenAI model when
that provider is configured on the org.
Using capability aliases
The cap:<name> syntax in the model field tells the gateway to pick a
concrete model at routing time based on the capability's stable
intent rather than a pinned model id. The gateway weighs your org's
capability overrides and current provider health when resolving which
model actually handles the request, so a cap:reason-heavy call can land
on a different concrete model over time without your code changing.
Five default capabilities resolve to a concrete model: reason-heavy,
tool-call-strict, long-context-128k, vision, and json-mode. A
sixth, cheap-fast, is accepted too but currently routes like auto
-- see capability aliases.
import os
import anthropic
client = anthropic.Anthropic(
base_url="https://gateway.iq-routing.com",
api_key=os.environ["IQ_API_KEY"],
)
message = client.messages.create(
model="cap:reason-heavy",
max_tokens=1024,
messages=[{"role": "user", "content": "Plan the migration in three phases."}],
)
The resolved provider plus model surface in the x-iq-routing
response header (a JSON-encoded payload with chosen_provider and
chosen_model fields, among others), so the SDK consumer can inspect which concrete
model handled the request. See the capability aliases
docs for the full list of default capabilities and
how to override them. You can override the default capability mapping
for your org via the capabilities dashboard editor at
/settings/capabilities.
Common gotchas
The Anthropic SDK's default max_retries is 2. The gateway's
429 already cooldowns; SDK retries on top double-bill the cooldown
window. Set max_retries=0.
stream=True works against the gateway. The SDK's
MessageStreamManager consumes the SSE just as it does against
api.anthropic.com. The X-Request-Id header is on the underlying HTTP
response, available via the event.message.id field on each streamed
chunk for correlation against the gateway's requests view.
Tool use (tools parameter) routes through the gateway transparently. The
gateway forwards the tool-use blocks in both directions without rewriting
the schema. Beta features (anthropic-beta header) are passed through: the
gateway forwards the header verbatim to the upstream provider and does not
evaluate it or reject the request over it. Every error body on this surface
carries a request_id field for correlation against the requests view.
The gateway's own worst case is up to 600 seconds per provider
attempt and 1800 seconds (30 minutes) across a full fallback chain
before it gives up with a 504. Most requests finish in a few seconds;
pass an explicit timeout on the client if your product needs to fail
faster than either ceiling.
When every upstream provider in the fallback chain is exhausted before
that ceiling is hit, the gateway returns 503
(rendered as an overloaded_error in the Anthropic error body). That
response does not carry a Retry-After header, so back off on your own
schedule. The one place the gateway does emit Retry-After is a 429
(the per-team rate-limit window); the Anthropic SDK does not auto-retry on
either, so a caller-side queue should read Retry-After on the 429 and
apply its own backoff on the 503.