IQ CLI
Self-serve availability: the
iqbinary shown throughout this page is not yet available for public download (see Install, below). Enterprise customers can reach it today through the self-hosted bundle.
The IQ CLI is the local terminal for the gateway. Running iq with no
arguments drops you into an interactive shell; everything else is a
one-shot subcommand:
iq/iq shell-- interactive REPL. Type a prompt, get a response with routing telemetry, repeat.iq prompt "..."-- single-prompt round-trip with a routing telemetry box (tier, model, focus mode, cache, latency split, cost).iq batch prompts.txt-- run a prompt file end-to-end with a per-prompt line and an aggregate table at the close.iq smoketest-- drive a canned routing-correctness corpus through the gateway and check which tier each prompt landed on.iq agent "..."-- a dry-run planner. Proposes a plan; executes nothing.iq generate-prompts-- emit a synthetic prompt file at the requested difficulty mix.
iq login and iq logout handle credentials -- see Sign in, below.
Install
The standalone iq CLI is not available for download yet. There
is no public release channel today. Enterprise customers can request
the self-hosted iq-gateway bundle, which includes the CLI as part of
its single-binary install -- see
Bundle mode (Enterprise) below, or contact
sales.
Sign in
iq login
iq login opens the dashboard's CLI access tokens page in your default
browser. Issue a token under Settings → CLI access tokens, paste it
back at the prompt. The CLI stores the token at
~/.config/iq-routing/credentials.toml (%APPDATA%\iq-routing\credentials.toml
on Windows) at mode 0600. Plaintext file storage with a restrictive file
mode is the mechanism today; OS-native credential storage (Keychain on
macOS, Credential Manager on Windows, Secret Service on Linux) is on the
roadmap.
CLI access tokens carry the prefix iqp_. The CLI also accepts an
org API key (prefix gw_live_) issued from the dashboard's Keys page,
so either credential class works. The gateway verifies each class
separately on the prefix, and a personal access token is scoped
narrower than an org-level API key.
For CI or any scripted run, skip iq login entirely: set IQ_TOKEN
to a CLI access token or an org API key in the environment and the
CLI reads it directly instead of the credentials file. IQ_GATEWAY_URL
points that same run at a different gateway (a self-hosted bundle, for
instance), overriding both the credentials file and the default. Both
are read fresh on every invocation, so nothing is written to disk.
Single-prompt mode
$ iq prompt "Summarise the last quarter's earnings call in three bullets"
Routing
Tier: medium
Complexity: 6.2 / 10
Chosen model: gpt-5.6-luna
Focus mode: balanced (no bias)
Cache: miss (new prompt)
Latency
Total: 1842 ms
Gateway-side: 34 ms
LLM-side: 1808 ms
Tokens
Prompt: 17
Completion: 142
Cost
This request: $0.000482
Response
─────────────────────────────────────────────────────────────────────────
- Revenue grew 12 percent YoY, driven by the AI gateway segment.
- Margins compressed 80 bps as inference costs outpaced price increases.
- Guidance trims FY26 EPS by 3 percent to reflect the launch ramp.
─────────────────────────────────────────────────────────────────────────
Flags:
--raw-- print only the response body. No telemetry box. Useful for piping into another tool.--json-- print the full response as JSON with the telemetry inlined underx_iq_routing.--model-- pin a routing alias (auto,cheap,default, orfrontier) or a concrete provider/model string.autois used when this flag is omitted.--timeout-- request timeout in seconds. Default 60.
On failure, iq prompt prints error: <message> to stderr and exits
non-zero: 2 for a credentials problem (not logged in, or a corrupt
credentials file), 1 for anything the gateway itself returned (a 4xx,
a 5xx, a timeout, or a transport error). A 429 falls into that same
exit-1 bucket -- iq prompt, iq batch, and iq agent do not retry
it automatically. Only iq smoketest (and :smoke in the shell)
waits out the suggested window and resumes on its own; see Smoketest,
below.
Interactive shell
Running iq with no arguments is the default entry point, and it is
identical to iq shell. Each line you type is sent through the
gateway as a new turn: the shell renders the same telemetry box as
iq prompt, then the response body, then prompts again. Turns are numbered (iq[1]>, iq[2]>, ...) and prior
turns re-render dimmed above the active one so a screen recording
shows which turn is live. Ctrl-D or Ctrl-C exits cleanly.
iq
# or start on a different routing alias:
iq shell --model frontier
$ iq
Connected to https://gateway.iq-routing.com. Model: auto
Type :help for commands, :exit to leave.
iq[1]> Summarise the last quarter's earnings call in three bullets
── turn 1 (14:02:07) ────────────────────────────────────────────
Routing
Tier: medium
Chosen model: gpt-5.6-luna
Cache: miss (new prompt)
...
Response
────────────────────────────────────────────────────────────────
- Revenue grew 12 percent YoY, driven by the AI gateway segment.
...
────────────────────────────────────────────────────────────────
iq[2]> What was the margin figure again?
── turn 2 (14:02:41) ────────────────────────────────────────────
Routing
Tier: simple
Chosen model: gpt-5.6-luna
Cache: miss (new prompt)
...
Response
...
iq[2]> :exit
Meta-commands (anything starting with :):
| Command | Effect |
|---|---|
| :help | List every meta-command. |
| :exit, :quit, :q | Leave the shell. |
| :key <api-key> | Use a different key for this session only. :key alone shows the masked current key; :key clear reverts to the credentials-file token. |
| :gateway <url> | Point this session at a different gateway. :gateway alone shows the current URL; :gateway clear reverts. |
| :model <alias> | Switch the routing alias mid-session (auto, default, cheap, frontier, or a literal provider/model). :model alone shows the current value. |
| :reset | Clear the screen and reset the turn counter to 1. Does not touch credentials. |
| :smoke [N] [--model alias] [--limit N] | Run the smoketest harness without leaving the shell. The positional N and --limit are interchangeable; passing both with different values is an error. |
| :testkey | Print a reminder of how to paste a real key. Does not itself change the active key. |
Dropping a file to batch
Paste or drag a .txt, .md, .prompt, or .prompts file path into
the prompt -- as the only token on the line, resolving to a file that
exists -- to dispatch every line in it as a batch, at concurrency 4,
without leaving the shell:
iq[3]> ~/Desktop/support-queue.txt
dropped file support-queue.txt (23 prompts, model=auto, concurrency=4, 6 match canonical corpus)
1/23 ✓ exp=- obs=cheap gpt-5-nano cache-miss 0.61s $0.00009
complexity=simple (1.8/10) focus=balanced tokens=22→64 req=a1b2c3d4
> "How do I reset my password?"
...
Aggregate
Total prompts: 23
...
Your per-team rate limit depends on your plan: 240 requests per minute
on Free, 600 on Team, and 6,000 or more on Enterprise. A separate,
tighter burst ceiling applies per second, so a highly concurrent run can
see a 429 while still comfortably under the per-minute number. Check
your own limits on the rate-limits page in the dashboard.
The shell's own warning is a fixed heuristic rather than a reading of
your plan: it warns above 60 prompts in a single drop, which is the
most conservative tier's safe margin. Treat it as a nudge to check your
limit, not as the limit itself. Re-dropping the same file replays every
row from cache in a fraction of the original time.
exp= shows an expected tier only for prompts that match the built-in
smoketest corpus verbatim; anything else shows exp=- and only the
observed tier.
Batch mode
$ iq batch prompts.txt --concurrency=4 --out=results.jsonl
Reading prompts.txt: 20 prompts loaded
Concurrency: 4
1/20 ✓ simple gpt-5-nano cache-miss 0.84s $0.00012
2/20 ✓ medium gpt-5-mini cache-miss 1.92s $0.00048
3/20 ✓ simple gpt-5-nano cache-hit 0.04s $0.00000
...
20/20 ✓ complex opus-4-8 cache-miss 3.41s $0.00284
Aggregate
Total prompts: 20
Successes: 20
Failures: 0
Cache hit rate: 35%
Tier mix: simple 8 · medium 9 · complex 3
Latency p50: 1.42s
Latency p99: 4.18s
Total spend: $0.0341
Avg cost / prompt: $0.00171
Wrote results.jsonl with 20 rows.
Flags:
--concurrency=N-- in-flight requests. Defaults to 4. Caps at 16 (the safe per-org ceiling); higher values clamp with a stderr warning.--out=results.jsonl-- write one JSONL row per prompt with the full response body and the telemetry inlined underx_iq_routing.--model,--timeout-- same semantics asiq prompt.
The exit code goes non-zero when at least one prompt fails (any 4xx or
5xx response, a timeout, or a transport error). Failed prompts do not
abort the run; the aggregate reports successes and failures together.
A 429 counts as a failure here too -- iq batch does not auto-retry
a rate limit the way iq smoketest does, so lower --concurrency or
split the file if a run is tripping your plan's limit.
With --out, a failed prompt still gets a row: response and
x_iq_routing are null, and error holds the failure message.
Prompt-file format
One prompt per line. Blank lines and #-prefixed lines ignored.
# easy
What is the capital of France?
Translate "good morning" to Spanish.
# medium
Given this CSV row, output a JSON object with the columns as keys: 1,Alice,30
Prompt generator
iq generate-prompts --difficulty=mixed --count=20 --out=prompts.txt
Difficulty bands:
easy-- single-sentence factual lookups, single-step instructions. Target tiersimple.medium-- multi-step reasoning, structured output, short code snippets. Target tiermedium.hard-- long-context, ambiguous instructions, cross-domain synthesis. Target tiercomplex.mixed-- 40% easy, 45% medium, 15% hard. A reasonable default for exercising all three tiers in one file.
Counts cap at 1000 per file. The generator does not call the gateway; it is a pure offline emitter.
Smoketest
iq smoketest
iq smoketest --limit 10 --model frontier
iq smoketest drives a canned corpus of 100-plus prompts through the
gateway and checks that each one landed in its expected tier
(cheap, default, or frontier). It is a routing-correctness check, not
a synthetic dry run: every prompt is a real request against your
provider keys, so a full run has a real, if small, cost. Prompts
dispatch sequentially, so the full corpus takes several minutes.
$ iq smoketest --limit 3
iq smoketest 3 prompts gateway=https://gateway.iq-routing.com model=auto
PASS sp-cheap-01 expected=cheap model=gpt-5-nano 84ms
> What is the capital of France?
PASS sp-default-01 expected=default model=gpt-5.6-luna 1204ms
> Write a Python function that takes a list of integers and returns the second largest. Handle ties and empty lists.
FAIL sp-frontier-01 expected=frontier model=gpt-5-mini 932ms (routed to 'default' tier; expected 'frontier' (model: gpt-5-mini))
> Design the architecture for a multi-tenant SaaS analytics product handling 50 million events per day. Cover ingestion...
Routing distribution (expected tier x observed tier)
...
totals: 2 pass 1 fail 3 distinct chosen-model values runtime 14.2s
Flags:
--limit N/-n N-- cap the run at the first N prompts (a head-of-list subset) instead of driving the full corpus.--model/-m-- routing alias for the whole run:auto,cheap,default, orfrontier, or a literalprovider/modelstring. Same alias semantics asiq prompt --model;autois used when this flag is omitted.--timeout-- per-request timeout in seconds. Default 60.
The run fails (exit code 1) if any prompt's response is empty, missing
telemetry, missing a chosen model, non-positive latency, missing cost
on a non-cache-hit, or routed outside its expected tier -- or
if the whole run produced fewer than four distinct chosen-model
values. A 429 mid-run waits for the suggested retry window once and
resumes rather than failing the prompt outright, since a full run can
brush against your plan's rate limit. The X-RateLimit-Scope header on
the 429 names which ceiling fired, and Retry-After gives the wait.
:smoke [N] [--model alias] [--limit N] runs the same harness inside
the interactive shell without spawning a separate process; see
Interactive shell, above.
Agent (dry-run planner)
iq agent "Find every TODO comment in the src/ directory and summarise them"
iq agent "Check whether api.example.com is up" --tools http_fetch
iq agent is a single-turn, dry-run planner. It sends your
instruction plus a small catalog of advisory tool schemas
(read_file, http_fetch) to the gateway and prints the model's
proposed plan: either a numbered list of tool calls or a prose
explanation.
The agent never executes anything. No file is read, no URL is fetched, no code runs. The tool schemas exist so the model can propose concrete, well-formed steps -- the operator reads the plan and runs whatever looks right by hand. Do not build automation on top of this command expecting the proposed steps to have actually happened; nothing downstream of the plan is real until you make it real.
$ iq agent "Find every TODO comment in the src/ directory and summarise them"
iq agent (planner / dry-run) session=iqa_4f9c2e1a08b7c391 model=auto tools=http_fetch,read_file
instruction: Find every TODO comment in the src/ directory and summarise them
Routing telemetry
model=gpt-5-mini cache-miss 0.81s $0.00031 complexity=medium (5.1/10) tokens=41→96
Proposed plan (advisory only -- nothing executed)
1. read_file(path='src/index.ts')
2. read_file(path='src/utils.ts')
Model commentary:
Run these two reads, then grep the results for "TODO" and summarise
each hit with its surrounding context.
Run any of these steps yourself if they look right; the agent never
executes tools.
Flags:
--model/-m-- routing alias or literalprovider/model. Same semantics asiq prompt --model.--tools-- comma-separated list of tool names the planner may propose. Built-ins:read_file,http_fetch. Defaults to both.--theme--light(default) ordark.
Every turn carries an X-IQ-Session-Id header
(iqa_<16-hex-chars>), so the dashboard's Sessions view rolls the
telemetry under one row even though the agent itself never loops.
Tokens, scopes, and revoke
The CLI authenticates with a personal access token issued from Settings → CLI access tokens. Default scopes:
read:routing-- view routing decisions.write:prompt-- send prompts via the CLI.
Wider scopes (read:audit, write:settings) are opt-in at issue time.
Tokens display the plaintext exactly once at issue. The dashboard
shows only the prefix on subsequent reads.
iq logout removes the local credentials file. To revoke server-side,
open Settings → CLI access tokens and click Revoke on the token row.
What bounds a token's spend
A personal access token is bounded at the organisation level, not at the credential level. Your org's daily and monthly budgets both apply to CLI traffic, and the daily budget is enforced before the call is made rather than after, so a token cannot spend past it. Per-token controls are a different story: unlike an org API key, a personal access token carries no per-credential spend ceiling and no per-credential rate limit of its own, so it inherits your team and plan defaults instead.
The practical effect is that a token is a narrower credential than an API key in what it can reach, and a blunter one in how finely you can cap it. If you need a hard per-credential ceiling, issue an org API key scoped to a team and set the ceiling there.
Bundle mode (Enterprise)
There is no self-hosted distribution on the self-serve plans. The
standalone gateway bundle is the Enterprise on-prem option, arranged
with us rather than downloaded, and it ships the CLI as part of the
single-binary install. iq-gateway prompt "..." posts to the local
gateway at http://localhost:8000; the same telemetry box renders.
iq-gateway serve # boots the gateway + opens the dashboard
iq-gateway prompt "..." # uses the local gateway
Some of the bundle's deeper analytics surfaces render a Limited mode
banner: the on-prem install doesn't run the same aggregate reporting
queries as the hosted product. Routing, caching, and per-request
telemetry all run end to end in the bundle; only the aggregate
analytics views are reduced.
Troubleshooting
Token in {path} is not a recognised IQ credential-- the saved value carries no known prefix. Re-runiq loginand paste an org API key (gw_live_...) or a CLI access token (iqp_...).Credentials file is group- or world-readable-- POSIX safety check. Runchmod 0600 ~/.config/iq-routing/credentials.toml.Gateway returned 401-- the token expired or was revoked. Runiq loginto reissue.- Telemetry box shows
telemetry unavailable-- the gateway you are pointed at is older than the telemetry header, or telemetry has been turned off for that deployment. Update to a current gateway version, or re-enable the telemetry header on the deployment you are calling.
See also
/docs/quickstart-- SDK integration via OpenAI/Anthropic clients./docs/api-reference-- REST surface details./integrate-- drop-in code snippets and aUse the IQ CLIpanel.