Skip to content
IQ Routing

Routing

When a request arrives with model="auto", the gateway assigns it two labels: a complexity tier (simple, medium, complex) and a thinking budget tier (none, low, medium, high). Both labels ride on the request row and on the x-iq-routing telemetry header so you can see what the gateway decided for any request.

Reading the routing decision

curl -i https://gateway.iq-routing.com/v1/chat/completions \
  -H "Authorization: Bearer gw_live_xxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "messages": [{"role": "user", "content": "Summarise this support ticket in two sentences."}]
  }'

The x-iq-routing response header is a JSON-encoded payload carrying the chosen provider and model, the complexity tier, the thinking budget, and the cache-hit flag for that request. X-Request-Id is the correlation handle; open /requests/<uuid> in the dashboard for the full decision trace, including any fallback hops.

Picking the provider

The complexity tier maps to a model tier: simpler prompts route to cheaper, faster models and harder prompts route to your chosen flagship, configurable per team. A preferred-model list overrides the default when you set one, and an allowed-model list bounds the search so a runaway prompt cannot escalate to the most expensive tier.

Within one conversation, routing takes your provider's prompt cache into account. A later turn stays on the model that served the previous turn when that model still fits the routing option for the new request and reusing its cached prompt is expected to make it the cheaper choice; otherwise, the turn is routed as a fresh one. This applies to auto, optimal, cheap, frontier, and custom; named models are unaffected, since a named model already stays with its provider. For custom and optimal, a conversation stays only on a model the alias would already pick from -- stickiness never brings in a model from outside that set. On auto, it stays only on a model in the same tier as the one auto would pick, or a higher one. On cheap, it stays only when that is expected to make the turn cheaper than cheap's own choice, and never on your top-tier flagship. On frontier, it stays only on a model that meets frontier's quality level for the request, so a conversation can move to a more capable model from another connected provider when a request needs it. A turn the previous model can't serve, such as one longer than its context window, is routed as a fresh one. Send the same X-IQ-Session-Id header (also used for agent sessions) on every request in the conversation for reliable stickiness. Agent clients that send their own session id, such as Claude Code or Codex, also work. The link lapses after 30 minutes without activity, after which the next request routes fresh. If that model fails mid-conversation, that request falls back normally.

Routing aliases

auto is one of nine routing aliases you can pass as model:

  • auto -- picks a model for each request from the providers your org has connected, balancing quality and cost, and automatically reflects any provider key you add or remove.
  • optimal -- a curated, fixed selection of models that IQ Routing maintains centrally. If your org doesn't hold the key optimal's usual pick needs, IQ Routing serves the nearest model your org does hold a key for instead of failing.
  • cheap -- routes across a cost-capped ladder that never reaches your top-tier flagship, and never costs more than auto for the same request. Like auto, it reflects your org's held provider keys rather than a fixed roster.
  • frontier -- aims for the most capable models your connected providers offer, serving stronger models as requests get harder. Like auto, it reflects your org's held provider keys rather than a fixed roster.
  • custom -- routes using your org's own routing table, which you build on the dashboard's Routing tab (Team plan and up). It can include models from any provider you hold a key for, plus any custom endpoint you've added. If your org hasn't saved a table, custom behaves like auto.
  • dynamic -- chooses a model separately for each request from the providers your organization has connected, based on what that request needs (such as tools, images or a long context), balancing quality and cost (Team plan and up).
  • auto-coding -- task-intent alias for coding workloads; resolves through the frontier ladder today.
  • auto-research -- task-intent alias for research workloads; resolves through the frontier ladder today.
  • default -- resolves to one model the gateway has chosen as a stable starting point; behaves like a concrete id from there, including the family lock described below.

auto-coding and auto-research are task-intent aliases: they name the kind of work instead of a cost/quality tier directly, and the gateway currently maps both onto the frontier ladder. Use them when you want the call site to say what the request is rather than which tier to route it through; use frontier directly when you want to name the tier.

A request that doesn't name a model uses your org's Routing mode setting, which is auto unless you change it -- previously this resolved to one fixed model instead. Routing mode offers six of the aliases above (auto, optimal, cheap, frontier, custom, dynamic); setting it to custom or dynamic needs the Team plan or above.

Customising the ladder behind an alias

On the Team plan and up you can customise the model ladder behind auto, cheap, and frontier, choosing which models sit in which complexity tier, through the admin API. Because auto-coding and auto-research resolve through the frontier ladder, a custom frontier ladder governs them too. That is how you keep auto pointed at the models you actually trust for hard work while holding cheap traffic on whatever you have decided is good enough, without changing a line of client code. Free stays on the stock ladder.

In the dashboard, the Routing tab (Team plan and up) lets you build one table instead: custom (see Routing aliases above) is a separate routing table that the custom alias routes on directly, rather than overriding what auto, cheap, or frontier each pick.

Separately, on the Alias overrides page (/settings/aliases) in dashboard settings, or through the admin API, default, auto, cheap, and frontier can each be pointed at one specific model instead -- available on every plan, including Free, unlike the ladder customisation above (optimal is fixed and doesn't take an override, and custom is already your own table). Only an organization's owners and admins can set or clear one of these; other members can see the current value but not change it, and it can only point at a model from a provider your organization has connected.

Fallback loop

If the chosen provider returns a retryable error (429, 500, 503, network timeout), the gateway falls back to the next model in that tier's chain. The chain is ordered and finite, each rung carries a per-call timeout, and the whole fallback loop runs under a hard wall-clock ceiling, so a provider that accepts a connection and then goes quiet cannot hold your request open indefinitely. The requests browser records the provider that actually served each request, whether a fallback was used, and the per-rung attempt trail, so a degrade reads as "asked for X, served Y" rather than a bare flag.

Sending a concrete model id

Pass an explicit model id (claude-opus-5, claude-sonnet-5, claude-opus-4-8, claude-haiku-4-5, gpt-5-mini) instead of auto when your client is bound to one provider's conventions. A concrete id is read as a family lock rather than as a pin: IQ Routing may serve a different model from the same provider to suit the request, but every candidate and every fallback rung is drawn from that provider's own ladder, so a coding harness built around one vendor's system prompt and tool protocol never lands on a foreign model mid-loop. The model can move up or down within that provider from turn to turn. This is the same behavior for every provider you can hold a key for, not only OpenAI, Anthropic, and Google -- naming a DeepSeek or Z.ai model, for example, family-locks it exactly the same way.

Every provider's ids carry a <prefix>/ form too (for example deepseek/deepseek-flash, anthropic/claude-sonnet-5); see Models for the full mapping. Anthropic and OpenAI ids also work unprefixed, as shown above.

That means the model you name is a family and a starting point, not a guarantee of the exact id. For deliberate steering, use one of the routing aliases above instead of a concrete id. auto, optimal, cheap, frontier, custom, auto-coding, and auto-research route across providers and are unaffected by the family lock. default is the exception: it resolves to one fixed model the gateway chooses, which then triggers this same family lock as if you had named that model directly -- unless your org holds no key for that provider, in which case the request falls open to auto-style routing instead.

These ids are pinned exactly rather than family-locked: naming gpt-6-astra, gpt-6.1-sol, gpt-6-sol, claude-opus-5-5, or claude-fable-5-1 (bare, or with its provider prefix) serves exactly that model whenever it can serve the request. If it can't -- the request needs a feature that model doesn't support, or the model is temporarily unavailable -- IQ Routing falls back to another model from the same provider, the same fallback you'd see from a family-locked id. When that happens, the x-iq-served-model response header names the model that actually answered, the same way it does whenever a named model is served by a different one. Every other model id on this page keeps the family-lock behavior above.

Routed requests can use GPT-6.1 Sol. Naming gpt-6-sol continues to request GPT-6 Sol. Your routing mode and organization permissions still apply. Check x-iq-served-model to see which model answered. See Models for GPT-6.1 Sol request settings and pricing.

Anthropic-format requests (for example from Claude Code, through /v1/messages) follow the same rule: sent with a routing alias (auto, optimal, cheap, or frontier), they may be served by any connected provider whose models support the request's features (tools, images, context length). Naming a concrete Claude model keeps the request with Anthropic, the same as any other named id.

claude-fable-5-1 is Anthropic's current Fable model. Naming it directly serves exactly that model whenever it can serve the request, per the exact pins above; routing aliases may also select it. An older id, claude-fable-5, is still accepted if you name it directly and is served on the Anthropic family ladder. See capability aliases for cap:orchestrate, the only way to reach claude-fable-5 itself, with its own limits and the data-retention requirement Anthropic places on it.