Capability aliases
A capability alias is a stable, intent-named handle you can write as the model, such as model=cap:reason-heavy, instead of pinning a specific model at every call site. Your organization chooses which model each capability uses, so moving an agent loop to a new model family is one dashboard change rather than a code change. Six default capabilities cover the most common kinds of agent work, and you can override any of them.
The six default capabilities
The gateway recognises six default capability names. Five of them
resolve to a concrete model and thinking budget that the gateway
maintains and tunes; you write the intent-named handle and the gateway
keeps the mapping current as the provider mix changes. The sixth,
cheap-fast, is a recognised name with no mapping today (see below).
You can override any mapped default with your own model pick per org
(see below).
A capability's thinking budget resolves to one of four settings, none,
low, medium, or high, where medium sits between a cheap nudge and
a full reasoning pass. Each setting maps onto whatever the destination
provider actually supports, so the same capability produces an Anthropic
thinking budget, an OpenAI reasoning effort, or a Gemini thinking budget
without the alias or your call site changing. If your org runs thinking
budgets in auto mode, the gateway picks the tier per
request instead of the capability pinning one, and every request records
whether the budget came from the client, the org default, or auto, so the
choice is auditable in the requests browser.
The reason-heavy capability covers plan-style steps in a multi-step
agent loop: anywhere the agent decomposes a goal into subgoals, and any
critique step that has to evaluate a prior agent output against a
specification. It resolves to a high-reasoning flagship with an extended
thinking budget.
The tool-call-strict capability covers the tool-arg-formatter step in an
agent loop, anywhere the agent emits a function call against a JSON
schema. It resolves to a model with strong strict-mode function calling so
the JSON shape lands reliably on a tight schema.
The long-context-128k capability covers a prompt that carries a large
attachment (a long meeting transcript, a multi-document research brief,
the full context window of a prior agent loop) where the response needs to
reason over the full attachment. It resolves to a cost-efficient
long-context model.
The vision capability covers any request that includes an image
attachment in the message array. It resolves to a cost-efficient
vision-capable model.
The cheap-fast capability is reserved for the high-volume, low-stakes
parts of an agent loop: text classification, sentiment scoring, and
short Q+A against a small context. It has no bound model today. The
gateway accepts cap:cheap-fast and routes it exactly as it routes
model: "auto", picking a model from the providers your org holds
keys for. A short classification prompt usually lands on your
provider's lowest-cost model, but nothing caps the pick, so a more
demanding prompt can still be served by a frontier model. Every
response to a cap:cheap-fast request carries an x-iq-served-model
header of
{"capability":"cheap-fast","cap_unresolved":true,"cap_unresolved_reason":"no_row"}.
The model that answered is named in the response body's model field
and, on non-streaming responses, in the x-iq-routing header's
chosen_model and chosen_provider fields.
The json-mode capability covers a strict-JSON response shape where the
input is a free-text prompt rather than a tool-call schema. It resolves to
a cost-efficient model that lands JSON mode reliably.
When to use which capability
Reason-heavy synthesis is the canonical fit for cap:reason-heavy. When
the agent has to combine multiple retrieval results into a coherent
answer, the synthesis step rewards the high-quality multi-step reasoning
path more than any other step in the loop. The thinking-budget=high
setting pays for itself on the synthesis step because the model has to
reconcile contradictions across the retrieved chunks and produce a
single output that respects all of them.
Tool-call execution is the canonical fit for cap:tool-call-strict.
When the agent has to emit a structured function call against a tight
JSON schema, the strict-mode function-calling path lands the schema
correctly more often than the alternatives. The provider pick on this
capability is load-bearing: swapping it for a weaker alternative
measurably raises the odds of a malformed tool call, which is exactly the
failure mode this capability exists to avoid. Reach for
cap:tool-call-strict any time a dropped or malformed schema would break
your agent loop, not only on your highest-volume tool-call step.
Long-context retrieval is the canonical fit for cap:long-context-128k.
When the agent is reading a 100k-token document and answering a question
against the full context, the cost-efficient long-context window beats
the alternatives by an order of magnitude on per-token cost without
giving up answer quality on the retrieval-style question shape.
Vision input is the canonical fit for cap:vision. When the agent is
reading a screenshot or chart and answering a question about it, the
vision-capable model with no thinking-budget overhead lands the right
output without paying the high-thinking premium on a perception-bound
task. Vision tasks rarely benefit from extended chain-of-thought because
the bottleneck is the visual parse, not the reasoning step.
Cost-sensitive simple QA is what cap:cheap-fast is reserved for, but
while it has no bound model it adds nothing to model: "auto". For a
classification or sentiment step today, send auto, or give
cheap-fast a fixed low-cost target with a per-org override (see
below) so the handle in your agent code stays stable.
JSON-mode structured output is the canonical fit for cap:json-mode.
When the agent has to emit a strict JSON schema from a free-text prompt
(not a tool-call schema), the schema-conformant JSON output path lands
the shape correctly without the overhead of the tool-call wrapper. Use
cap:tool-call-strict instead when the output is a function call; use
cap:json-mode when the output is a free-form JSON object that the
downstream code parses directly.
Which model a capability uses
A capability uses your organization's override when you've set one, and IQ's default otherwise. That model is served first, within your team's model allow-list. If it returns a retryable error, IQ falls back to another model automatically, and the x-iq-served-model response header names the model that actually served the request and why. Override changes reach live traffic within about a minute. Each request's detail records the capability, the model it used, and whether that came from your override or the default.
cap:orchestrate
cap:orchestrate is a seventh capability handle, structurally separate
from the six defaults above, and the only way to reach claude-fable-5,
the older Fable id. No routing alias selects it, and naming
claude-fable-5 directly is accepted but served by another Anthropic
model instead -- check the x-iq-served-model response header to see
which one. cap:orchestrate has its own limits, described below.
Anthropic serves claude-fable-5 only to organisations with standard
30-day data retention. If your own Anthropic organisation runs zero data
retention without Anthropic's express authorisation for this model,
Anthropic rejects the call and the request moves down the fallback tail,
so another model answers; check the response's served-model header to
see which one.
A safety refusal that comes back as a well-formed 200 is not treated as a
successful response on a cap:orchestrate route. The gateway detects the
refusal structurally and treats it as a routing failure, which triggers the
same automatic fallback described above: the request moves to the next model
in the fallback tail, and the x-iq-served-model header names whichever
model actually answered. This detection covers non-streaming requests only
in v1; a streaming cap:orchestrate call does not get the same structural
check, so a streamed response still needs your own client-side handling for
a refusal.
Two per-request limits apply on cap:orchestrate routes specifically, and
both are hard stops with an explicit error rather than a silent
substitution. A per-request spend ceiling, $2.50 by default, is computed
as a worst-case cost estimate before the request ever dispatches; a
request whose estimate exceeds the ceiling never reaches the model, and the
gateway returns a 4xx error naming both the estimate and the cap. A
tool-call ceiling, 8 tool calls by default, caps how many tool calls a
single response can carry; a response that would exceed it returns an
error instead of the tool-call fan-out. Both numbers are product limits
that can change, not a fixed contract -- treat the values above as the
current defaults rather than something to hardcode against.
claude-fable-5 is priced at $10 per million input tokens and $50 per
million output tokens, which is what the spend-ceiling estimate is
computed against.
Focus mode
Your organization's focus mode also applies to capabilities: cost reduction favors lower-cost models and lower thinking budgets, quality favors premium models and higher budgets, and balanced uses your mapping as written. It's the quickest way to shift an agent loop's cost without changing code.
Override the default mapping
The dashboard editor at /settings/capabilities
surfaces the per-org capability table. Each row carries the capability
name, the resolved (provider, model, thinking_budget) tuple, the row
source (global default when no override exists, per-org override when
the operator has written a row), and the priority column. Priority decides
which row wins when your org holds more than one override for the same
capability: the resolver always reads the highest-priority row and ignores
the rest, so writing a new override at a higher priority supersedes an
older one without deleting it first. Writes are clamped to the 100-999
range; the seeded global defaults sit at 100, the floor of that range.
The Edit drawer opens a per-row form with the same four fields; submitting the
form writes the override row and emits a capability_override_set audit
event.
cheap-fast has no global default row, so it does not appear in the
editor; to bind it, write an override with the API call below, naming
a model on a provider your org holds a key for.
A Test capability button on the Edit drawer lets the operator preview the resolved tuple plus the live cost estimate against a synthetic prompt before committing the override row, so a fat-fingered provider pick or a typo in the model name surfaces in the preview pane rather than landing on the live agent loop. The prompt is a short editable field on the drawer, so the test is one throwaway call of your own wording rather than a sample of your live traffic, and the result panel reports the chosen provider, the chosen model, the latency, the per-call cost, and the response text.
Every override write records an audit entry under the
capability_override_set action. The entry carries the actor
email, the previous tuple, the new tuple, and the request id of the
write so the operator can replay the change against the request browser.
The write also triggers a webhook delivery to every active subscriber
under the org's webhook subscription list, so a downstream automation
can react to the override (for example, posting a Slack notice to the
ops channel or invalidating a downstream cache).
curl -X POST https://gateway.iq-routing.com/admin/capabilities \
-H "Authorization: Bearer <session token>" \
-H "Content-Type: application/json" \
-d '{
"org_id": "11111111-2222-3333-4444-555555555555",
"capability": "reason-heavy",
"provider": "openai",
"model": "gpt-5",
"thinking_budget": "high",
"priority": 200
}'
The webhook event payload carries the same fields as the audit row plus the event name and the created-at timestamp:
{
"event": "capability_override_set",
"org_id": "...",
"capability": "reason-heavy",
"provider": "openai",
"model": "gpt-5",
"thinking_budget": "high",
"priority": 200,
"actor_email": "...",
"created_at": "..."
}
The cap:<name> syntax
The gateway accepts the cap: prefix on the request's model field in
both the OpenAI-compatible chat-completions surface and the Anthropic-
compatible messages surface. The prefix replaces the concrete model
name; the rest of the request shape is unchanged.
curl https://gateway.iq-routing.com/v1/chat/completions \
-H "Authorization: Bearer gw_live_xxxxxxxx" \
-H "Content-Type: application/json" \
-d '{
"model": "cap:reason-heavy",
"messages": [
{"role": "user", "content": "Decompose: ship a billing dashboard in two weeks."}
]
}'
The response carries the resolved concrete model on the model field of
the response body, so the agent loop can log the actual model that ran
without an extra round-trip to the gateway's admin surface. If automatic
fallback changed which model actually served the request, the
x-iq-served-model response header names that model, so the divergence
from the requested capability is visible from the response headers
alone.
The X-Request-Id header on the response correlates to the route-decision
row in the dashboard's /requests browser, where the resolved tuple
plus the focus-mode bias plus the capability source appear as columns on
the request detail page.