Skip to content
IQ Routing

Quickstart

New here? Start with Getting started to connect a provider, create an IQ API key, and check your first request.

Point the OpenAI Python SDK at the IQ Routing gateway. Three lines change.

from openai import OpenAI

client = OpenAI(
    base_url="https://gateway.iq-routing.com/v1",
    api_key="gw_live_xxxxxxxx",
)

response = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Summarise this contract."}],
)

The model="auto" ask picks a model from your connected providers, balancing quality and cost. Pinning a specific model (for example claude-haiku-4-5) keeps the request with that provider; name an exact-pin id (gpt-6-astra, gpt-6.1-sol, gpt-6-sol, claude-opus-5-5, claude-fable-5-1) for an exact model. See Models for GPT-6.1 Sol request settings and pricing.

Correlating gateway errors to requests

Every response carries an X-Request-Id header with the gateway's request id. The same id appears in the JSON body's request_id field on error responses. Search the requests browser at /requests for that id to see the full pipeline trace, including fallback hops and the provider's raw response.

curl -i https://gateway.iq-routing.com/v1/chat/completions \
  -H "Authorization: Bearer gw_live_xxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{"model":"auto","messages":[{"role":"user","content":"hi"}]}'

The response includes X-Request-Id: <uuid>. Open /requests/<uuid> to inspect the routing decision.

Anthropic SDK

The Anthropic SDK works the same way against /v1/messages:

from anthropic import Anthropic

client = Anthropic(
    base_url="https://gateway.iq-routing.com",
    api_key="gw_live_xxxxxxxx",
)

message = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=256,
    messages=[{"role": "user", "content": "Say hi."}],
)

Use the bare origin for the Anthropic base URL (no /v1); the SDK appends /v1/messages itself.

Responses API

The /v1/responses surface is stateful: IQ keeps the conversation, so a follow-up call passes previous_response_id instead of resending the transcript.

from openai import OpenAI

client = OpenAI(
    base_url="https://gateway.iq-routing.com/v1",
    api_key="gw_live_xxxxxxxx",
)

first = client.responses.create(
    model="auto",
    input="Summarise this contract.",
)

second = client.responses.create(
    model="auto",
    input="Now do the same for the addendum.",
    previous_response_id=first.id,
)

Each response carries an id. Pass it back as previous_response_id on the next call and the gateway continues that conversation.

Next

  • Auth -- API key scopes, rate limits, budgets, and revocation.
  • Routing -- the full set of model aliases auto supports and how to pin a ladder.
  • Caching -- how repeat prompts get served from cache and how to bypass it.
  • API reference -- request and response schemas for every endpoint.