Skip to content
IQ Routing

Codex CLI

Codex speaks the OpenAI Responses API, and so does the gateway: point a custom model provider at the gateway base URL, give it a gateway key, and every Codex turn (planning, edits, shell commands, streamed tool calls) routes through IQ. Codex reads provider configuration from ~/.codex/config.toml.

Drop-in switch

Add a custom model provider to ~/.codex/config.toml and make it the default:

model = "auto-coding"
model_provider = "iq-routing"

[model_providers.iq-routing]
name = "IQ Routing"
base_url = "https://gateway.iq-routing.com/v1"
env_key = "IQ_API_KEY"
wire_api = "responses"

Then export your gateway key under the name you gave env_key and run Codex:

export IQ_API_KEY="gw_live_xxxxxxxx"
codex

base_url keeps the /v1 suffix, since Codex calls /v1/responses. wire_api = "responses" pins the wire format Codex 0.142 and later actually speak; the gateway implements the stateful Responses API end to end, so streamed tool calls and previous_response_id continuation work the same as they do against OpenAI directly. model = "auto-coding" is a task-intent alias that resolves through the frontier ladder today, the same ladder behind the frontier alias; use auto to pick a model from your connected providers, or pin a concrete id or a capability alias instead (see below). Customizing your org's frontier ladder -- a Team-plan admin-API setting -- governs auto-coding traffic too. See routing aliases for the full picture.

Opt-in wrapper instead of a default

If you would rather keep plain codex on its existing login and route through the gateway only when you choose, skip the config-file default and pass the same provider as -c overrides from a shell function:

iqcodex() {
  codex \
    -c model_provider=iq \
    -c model=auto-coding \
    -c 'model_providers.iq.name="IQ Routing"' \
    -c 'model_providers.iq.base_url="https://gateway.iq-routing.com/v1"' \
    -c 'model_providers.iq.env_key="IQ_API_KEY"' \
    -c 'model_providers.iq.wire_api="responses"' \
    "$@"
}

iqcodex then routes through IQ while bare codex stays exactly as it was.

Verify it routes

Start Codex in a scratch repo and give it a one-line task (“list the files in this directory”). The tool call streams, executes, and the summary renders with the gateway's standard latency. Open your dashboard at /requests and each Codex turn is its own row, with the chosen model, the cost, and the cache-hit status. The response also carries the x-iq-routing header (chosen_provider, chosen_model, cache_hit), so you can confirm which concrete model handled the turn, plus an X-Request-Id header you can paste into /requests/<uuid> to pull up that exact turn. An invalid key returns 401; a revoked key (from /keys) returns 403 API key revoked. No row after a turn that otherwise looked normal usually means model_provider stayed on its default -- see Common gotchas below.

Using capability aliases

Set model to a cap:<name> alias to route by stable intent instead of a pinned model id. Five default capabilities resolve to a concrete model: reason-heavy, tool-call-strict, long-context-128k, vision, and json-mode. A sixth, cheap-fast, is accepted too but currently routes like auto -- see capability aliases:

model = "cap:reason-heavy"
model_provider = "iq-routing"

Codex sends that string through as the model field on each call. Resolution checks your org's capability table first, then the global default, and folds in your account's focus-mode setting; if the resolved model can't be reached, the gateway falls back down the routing ladder instead of failing the call. See the capability aliases docs for the full set of defaults, the focus-mode interaction, and how to override them.

Common gotchas

Codex has no environment-variable override for its endpoint -- OPENAI_BASE_URL does nothing; the binary does not read it. The model_provider you select, whether the config.toml default above or a -c model_provider=... override, is the only thing that decides where a turn goes. If Codex is signed in with codex login (a ChatGPT subscription) and model_provider is left at its default, turns go straight to OpenAI on that subscription, never reach the gateway, and leave no row in /requests. Set model_provider to your iq-routing (or iq) block, as the config.toml default or per invocation with the wrapper above, to guarantee routing.

If you are following an older recipe that pins wire_api = "chat", drop it: recent Codex releases default to the Responses wire, and some no longer accept wire_api = "chat" at all, so a stale pin carried over from an old config can break the connection outright. Run codex --version and check which wire formats your installed build supports if a config that used to work stops connecting. The gateway itself still accepts both wire formats, so other OpenAI-shape clients that speak Chat Completions are unaffected.

Long agent sessions can spend fast, especially with shell-command loops. If a turn returns 429, check the X-RateLimit-Scope header to see which cap fired: rps, rpm, and tpm carry a Retry-After window to back off against (rps is a one-second window, rpm/tpm are sixty-second), while daily_budget, monthly_budget, and key_spend_ceiling have no rolling reset and need a manual budget increase. Watch your /dashboard spend bar during long runs. See the FAQ for the full list of scopes.

Function tools round-trip fully, streamed or not, including namespaced tool groups (Codex sub-agents and MCP-app tools): a dot-qualified ns.fn call still reaches the right sub-agent or MCP tool on the other side. Provider-hosted tools with no function body, such as web_search and image_generation, are the one category that does not cross yet -- the gateway drops those entries from the roster instead of forwarding them, so a Codex run that reaches for one falls back to its function tools rather than erroring the whole call.

On startup Codex refreshes model metadata from /v1/models. The gateway's response carries the top-level models array Codex reads alongside the standard OpenAI data array, so the refresh parses cleanly and routing is unaffected.

Codex sends its own system prompt and tool definitions; the gateway passes the prompt through untouched and routes on the full request. Each Codex turn is its own row in /requests; the gateway does not collapse a multi-turn session into one entry.