Skip to content
IQ Routing

Anthropic · Python

The Anthropic Python SDK accepts a base_url argument on the client. Point it at the gateway and the standard client.messages.create calls route through. This recipe pins claude-haiku-4-5-20251001 directly, which keeps every call in the example with Anthropic; name claude-opus-5-5 or claude-fable-5-1 instead for an exact model. See the Using capability aliases section below, or pass model="auto", to let the gateway pick the model per request instead.

Drop-in switch

client = anthropic.Anthropic(
    base_url="https://gateway.iq-routing.com",
    api_key="gw_live_xxxxxxxx",
)

The gateway exposes the Anthropic surface at /v1/messages (no /v1 prefix in the SDK's internal routing; the SDK appends /v1/messages itself). Pass the gateway origin without /v1.

Verify it routes

import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://gateway.iq-routing.com",
    api_key=os.environ["IQ_API_KEY"],
    timeout=60.0,
    max_retries=0,
)

response = client.messages.with_raw_response.create(
    model="claude-haiku-4-5-20251001",
    max_tokens=64,
    messages=[{"role": "user", "content": "Summarise: routing-aware gateway."}],
)

print("X-Request-Id:", response.headers.get("x-request-id"))
parsed = response.parse()
print("Stop reason:", parsed.stop_reason)
print(parsed.content[0].text)

Drop the script into a fresh venv, pip install anthropic, set IQ_API_KEY in your environment, run. The X-Request-Id header is the gateway's correlation handle.

The model pin (claude-haiku-4-5-20251001) keeps every call with Anthropic; name claude-opus-5-5 or claude-fable-5-1 instead for an exact model. If you pass model="auto" the gateway picks a model from your connected providers instead, balancing quality and cost. The auto pick is cross-provider, so a turn can land on an OpenAI model when that provider is configured on the org.

Using capability aliases

The cap:<name> syntax in the model field tells the gateway to pick a concrete model at routing time based on the capability's stable intent rather than a pinned model id. The gateway weighs your org's capability overrides and current provider health when resolving which model actually handles the request, so a cap:reason-heavy call can land on a different concrete model over time without your code changing. Five default capabilities resolve to a concrete model: reason-heavy, tool-call-strict, long-context-128k, vision, and json-mode. A sixth, cheap-fast, is accepted too but currently routes like auto -- see capability aliases.

import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://gateway.iq-routing.com",
    api_key=os.environ["IQ_API_KEY"],
)

message = client.messages.create(
    model="cap:reason-heavy",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Plan the migration in three phases."}],
)

The resolved provider plus model surface in the x-iq-routing response header (a JSON-encoded payload with chosen_provider and chosen_model fields, among others), so the SDK consumer can inspect which concrete model handled the request. See the capability aliases docs for the full list of default capabilities and how to override them. You can override the default capability mapping for your org via the capabilities dashboard editor at /settings/capabilities.

Common gotchas

The Anthropic SDK's default max_retries is 2. The gateway's 429 already cooldowns; SDK retries on top double-bill the cooldown window. Set max_retries=0.

stream=True works against the gateway. The SDK's MessageStreamManager consumes the SSE just as it does against api.anthropic.com. The X-Request-Id header is on the underlying HTTP response, available via the event.message.id field on each streamed chunk for correlation against the gateway's requests view.

Tool use (tools parameter) routes through the gateway transparently. The gateway forwards the tool-use blocks in both directions without rewriting the schema. Beta features (anthropic-beta header) are passed through: the gateway forwards the header verbatim to the upstream provider and does not evaluate it or reject the request over it. Every error body on this surface carries a request_id field for correlation against the requests view.

The gateway's own worst case is up to 600 seconds per provider attempt and 1800 seconds (30 minutes) across a full fallback chain before it gives up with a 504. Most requests finish in a few seconds; pass an explicit timeout on the client if your product needs to fail faster than either ceiling.

When every upstream provider in the fallback chain is exhausted before that ceiling is hit, the gateway returns 503 (rendered as an overloaded_error in the Anthropic error body). That response does not carry a Retry-After header, so back off on your own schedule. The one place the gateway does emit Retry-After is a 429 (the per-team rate-limit window); the Anthropic SDK does not auto-retry on either, so a caller-side queue should read Retry-After on the 429 and apply its own backoff on the 503.

Try this in the playground

Run this snippet against your gateway key →