Routing
When a request arrives with model="auto", the gateway assigns it two
labels: a complexity tier (simple, medium,
complex) and a thinking budget tier (none, low, medium, high).
Both labels ride on the request row and on the x-iq-routing telemetry
header so you can see what the gateway decided for any request.
Reading the routing decision
curl -i https://gateway.iq-routing.com/v1/chat/completions \
-H "Authorization: Bearer gw_live_xxxxxxxx" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "Summarise this support ticket in two sentences."}]
}'
The x-iq-routing response header is a JSON-encoded payload carrying the
chosen provider and model, the complexity tier, the thinking budget, and
the cache-hit flag for that request. X-Request-Id is the correlation
handle; open /requests/<uuid> in the dashboard for the full decision
trace, including any fallback hops.
Picking the provider
The complexity tier maps to a model tier: simpler prompts route to cheaper, faster models and harder prompts route to your chosen flagship, configurable per team. A preferred-model list overrides the default when you set one, and an allowed-model list bounds the search so a runaway prompt cannot escalate to the most expensive tier.
Within one conversation, routing takes your provider's prompt cache
into account. A later turn stays on the model that served the previous
turn when that model still fits the routing option for the new request
and reusing its cached prompt is expected to make it the cheaper
choice; otherwise, the turn is routed as a fresh one. This applies to
auto, optimal, cheap, frontier, and custom; named models are
unaffected, since a named model already stays with its provider. For
custom and optimal, a conversation stays only on a model the alias
would already pick from -- stickiness never brings in a model from
outside that set. On auto, it stays only on a model in the same tier
as the one auto would pick, or a higher one. On cheap, it stays
only when that is expected to make the turn cheaper than cheap's own
choice, and never on your top-tier flagship. On frontier, it stays
only on a model that meets frontier's quality level for the request,
so a conversation can move to a more capable model from another
connected provider when a request needs it. A turn the previous model
can't serve, such as one longer than its context window, is routed as
a fresh one. Send the same X-IQ-Session-Id header (also used for
agent sessions) on every request in the conversation
for reliable stickiness. Agent clients that send their own session id,
such as Claude Code or Codex, also work. The link lapses after 30
minutes without activity, after which the next request routes fresh.
If that model fails mid-conversation, that request falls back normally.
Routing aliases
auto is one of nine routing aliases you can pass as model:
auto-- picks a model for each request from the providers your org has connected, balancing quality and cost, and automatically reflects any provider key you add or remove.optimal-- a curated, fixed selection of models that IQ Routing maintains centrally. If your org doesn't hold the keyoptimal's usual pick needs, IQ Routing serves the nearest model your org does hold a key for instead of failing.cheap-- routes across a cost-capped ladder that never reaches your top-tier flagship, and never costs more thanautofor the same request. Likeauto, it reflects your org's held provider keys rather than a fixed roster.frontier-- aims for the most capable models your connected providers offer, serving stronger models as requests get harder. Likeauto, it reflects your org's held provider keys rather than a fixed roster.custom-- routes using your org's own routing table, which you build on the dashboard's Routing tab (Team plan and up). It can include models from any provider you hold a key for, plus any custom endpoint you've added. If your org hasn't saved a table,custombehaves likeauto.dynamic-- chooses a model separately for each request from the providers your organization has connected, based on what that request needs (such as tools, images or a long context), balancing quality and cost (Team plan and up).auto-coding-- task-intent alias for coding workloads; resolves through the frontier ladder today.auto-research-- task-intent alias for research workloads; resolves through the frontier ladder today.default-- resolves to one model the gateway has chosen as a stable starting point; behaves like a concrete id from there, including the family lock described below.
auto-coding and auto-research are task-intent aliases: they name the
kind of work instead of a cost/quality tier directly, and the gateway
currently maps both onto the frontier ladder. Use them when you want the
call site to say what the request is rather than which tier to route it
through; use frontier directly when you want to name the tier.
A request that doesn't name a model uses your org's Routing mode
setting, which is auto unless you change it -- previously this
resolved to one fixed model instead. Routing mode offers six of the
aliases above (auto, optimal, cheap, frontier, custom,
dynamic); setting it to custom or dynamic needs the Team plan or above.
Customising the ladder behind an alias
On the Team plan and up you can customise the model ladder behind
auto, cheap, and frontier, choosing which models sit in which
complexity tier, through the admin API. Because auto-coding and
auto-research resolve through the frontier ladder, a custom
frontier ladder governs them too. That is how you keep auto
pointed at the models you actually trust for hard work while holding
cheap traffic on whatever you have decided is good enough, without
changing a line of client code. Free stays on the stock ladder.
In the dashboard, the Routing tab (Team plan and up) lets you build
one table instead: custom (see Routing aliases above) is a separate
routing table that the custom alias routes on directly, rather than
overriding what auto, cheap, or frontier each pick.
Separately, on the Alias overrides page (/settings/aliases) in
dashboard settings, or through the admin API, default, auto,
cheap, and frontier can each be
pointed at one specific model instead -- available on every plan,
including Free, unlike the ladder customisation above (optimal is
fixed and doesn't take an override, and custom is already your own
table). Only an organization's owners and admins can set or clear one
of these; other members can see the current value but not change it,
and it can only point at a model from a provider your organization has
connected.
Fallback loop
If the chosen provider returns a retryable error (429, 500, 503,
network timeout), the gateway falls back to the next model in that tier's
chain. The chain is
ordered and finite, each rung carries a per-call timeout, and the whole
fallback loop runs under a hard wall-clock ceiling, so a provider that
accepts a connection and then goes quiet cannot hold your request open
indefinitely. The requests browser records the provider that actually served
each request, whether a fallback was used, and the per-rung attempt trail, so
a degrade reads as "asked for X, served Y" rather than a bare flag.
Sending a concrete model id
Pass an explicit model id (claude-opus-5, claude-sonnet-5, claude-opus-4-8, claude-haiku-4-5, gpt-5-mini)
instead of auto when your client is bound to one provider's conventions. A concrete id is read as a family lock rather
than as a pin: IQ Routing may serve a different model from the same
provider to suit the request, but every candidate and every fallback rung
is drawn from that provider's own ladder, so a coding harness built
around one vendor's system prompt and tool protocol never lands on a
foreign model mid-loop. The model can move up or down within that
provider from turn to turn. This is the same behavior for every provider
you can hold a key for, not only OpenAI, Anthropic, and Google -- naming
a DeepSeek or Z.ai model, for example, family-locks it exactly the same
way.
Every provider's ids carry a <prefix>/ form too (for example
deepseek/deepseek-flash, anthropic/claude-sonnet-5); see
Models for the full mapping. Anthropic and OpenAI ids
also work unprefixed, as shown above.
That means the model you name is a family and a starting point, not a
guarantee of the exact id. For deliberate steering, use one of the routing
aliases above instead of a concrete id. auto, optimal, cheap,
frontier, custom, auto-coding, and auto-research route across
providers and are unaffected by the family lock. default is the
exception: it resolves to one fixed model the gateway chooses, which then
triggers this same family lock as if you had named that model directly --
unless your org holds no key for that provider, in which case the request
falls open to auto-style routing instead.
These ids are pinned exactly rather than family-locked: naming
gpt-6-astra, gpt-6.1-sol, gpt-6-sol, claude-opus-5-5, or claude-fable-5-1
(bare, or with its provider prefix) serves exactly that model whenever
it can serve the request. If it can't -- the request needs a feature
that model doesn't support, or the model is temporarily unavailable --
IQ Routing falls back to another model from the same provider, the same
fallback you'd see from a family-locked id. When that happens, the
x-iq-served-model response header names the model that actually
answered, the same way it does whenever a named model is served by a
different one. Every other model id on this page keeps the family-lock
behavior above.
Routed requests can use GPT-6.1 Sol. Naming gpt-6-sol continues to
request GPT-6 Sol. Your routing mode and organization permissions still
apply. Check x-iq-served-model to see which model answered. See
Models for GPT-6.1 Sol request settings and pricing.
Anthropic-format requests (for example from Claude Code, through
/v1/messages) follow the same rule: sent with a routing alias (auto,
optimal, cheap, or frontier), they may be served by any connected
provider whose models support the request's features (tools, images,
context length). Naming a concrete Claude model keeps the request with
Anthropic, the same as any other named id.
claude-fable-5-1 is Anthropic's current Fable model. Naming it
directly serves exactly that model whenever it can serve the request,
per the exact pins above; routing aliases may also select it.
An older id, claude-fable-5, is still accepted if you name it
directly and is served on the Anthropic family ladder. See
capability aliases for cap:orchestrate, the
only way to reach claude-fable-5 itself, with its own limits and the
data-retention requirement Anthropic places on it.