Skip to content
IQ Routing

IQ CLI

Self-serve availability: the iq binary shown throughout this page is not yet available for public download (see Install, below). Enterprise customers can reach it today through the self-hosted bundle.

The IQ CLI is the local terminal for the gateway. Running iq with no arguments drops you into an interactive shell; everything else is a one-shot subcommand:

  • iq / iq shell -- interactive REPL. Type a prompt, get a response with routing telemetry, repeat.
  • iq prompt "..." -- single-prompt round-trip with a routing telemetry box (tier, model, focus mode, cache, latency split, cost).
  • iq batch prompts.txt -- run a prompt file end-to-end with a per-prompt line and an aggregate table at the close.
  • iq smoketest -- drive a canned routing-correctness corpus through the gateway and check which tier each prompt landed on.
  • iq agent "..." -- a dry-run planner. Proposes a plan; executes nothing.
  • iq generate-prompts -- emit a synthetic prompt file at the requested difficulty mix.

iq login and iq logout handle credentials -- see Sign in, below.

Install

The standalone iq CLI is not available for download yet. There is no public release channel today. Enterprise customers can request the self-hosted iq-gateway bundle, which includes the CLI as part of its single-binary install -- see Bundle mode (Enterprise) below, or contact sales.

Sign in

iq login

iq login opens the dashboard's CLI access tokens page in your default browser. Issue a token under Settings → CLI access tokens, paste it back at the prompt. The CLI stores the token at ~/.config/iq-routing/credentials.toml (%APPDATA%\iq-routing\credentials.toml on Windows) at mode 0600. Plaintext file storage with a restrictive file mode is the mechanism today; OS-native credential storage (Keychain on macOS, Credential Manager on Windows, Secret Service on Linux) is on the roadmap.

CLI access tokens carry the prefix iqp_. The CLI also accepts an org API key (prefix gw_live_) issued from the dashboard's Keys page, so either credential class works. The gateway verifies each class separately on the prefix, and a personal access token is scoped narrower than an org-level API key.

For CI or any scripted run, skip iq login entirely: set IQ_TOKEN to a CLI access token or an org API key in the environment and the CLI reads it directly instead of the credentials file. IQ_GATEWAY_URL points that same run at a different gateway (a self-hosted bundle, for instance), overriding both the credentials file and the default. Both are read fresh on every invocation, so nothing is written to disk.

Single-prompt mode

$ iq prompt "Summarise the last quarter's earnings call in three bullets"

Routing
  Tier:           medium
  Complexity:     6.2 / 10
  Chosen model:   gpt-5.6-luna
  Focus mode:     balanced (no bias)
  Cache:          miss (new prompt)

Latency
  Total:          1842 ms
  Gateway-side:    34 ms
  LLM-side:      1808 ms

Tokens
  Prompt:           17
  Completion:      142

Cost
  This request:   $0.000482

Response
─────────────────────────────────────────────────────────────────────────
- Revenue grew 12 percent YoY, driven by the AI gateway segment.
- Margins compressed 80 bps as inference costs outpaced price increases.
- Guidance trims FY26 EPS by 3 percent to reflect the launch ramp.
─────────────────────────────────────────────────────────────────────────

Flags:

  • --raw -- print only the response body. No telemetry box. Useful for piping into another tool.
  • --json -- print the full response as JSON with the telemetry inlined under x_iq_routing.
  • --model -- pin a routing alias (auto, cheap, default, or frontier) or a concrete provider/model string. auto is used when this flag is omitted.
  • --timeout -- request timeout in seconds. Default 60.

On failure, iq prompt prints error: <message> to stderr and exits non-zero: 2 for a credentials problem (not logged in, or a corrupt credentials file), 1 for anything the gateway itself returned (a 4xx, a 5xx, a timeout, or a transport error). A 429 falls into that same exit-1 bucket -- iq prompt, iq batch, and iq agent do not retry it automatically. Only iq smoketest (and :smoke in the shell) waits out the suggested window and resumes on its own; see Smoketest, below.

Interactive shell

Running iq with no arguments is the default entry point, and it is identical to iq shell. Each line you type is sent through the gateway as a new turn: the shell renders the same telemetry box as iq prompt, then the response body, then prompts again. Turns are numbered (iq[1]>, iq[2]>, ...) and prior turns re-render dimmed above the active one so a screen recording shows which turn is live. Ctrl-D or Ctrl-C exits cleanly.

iq
# or start on a different routing alias:
iq shell --model frontier
$ iq
Connected to https://gateway.iq-routing.com.  Model: auto
Type :help for commands, :exit to leave.

iq[1]> Summarise the last quarter's earnings call in three bullets
── turn 1 (14:02:07) ────────────────────────────────────────────
Routing
  Tier:           medium
  Chosen model:   gpt-5.6-luna
  Cache:          miss (new prompt)
...
Response
────────────────────────────────────────────────────────────────
- Revenue grew 12 percent YoY, driven by the AI gateway segment.
...
────────────────────────────────────────────────────────────────

iq[2]> What was the margin figure again?
── turn 2 (14:02:41) ────────────────────────────────────────────
Routing
  Tier:           simple
  Chosen model:   gpt-5.6-luna
  Cache:          miss (new prompt)
...
Response
...

iq[2]> :exit

Meta-commands (anything starting with :):

| Command | Effect | |---|---| | :help | List every meta-command. | | :exit, :quit, :q | Leave the shell. | | :key <api-key> | Use a different key for this session only. :key alone shows the masked current key; :key clear reverts to the credentials-file token. | | :gateway <url> | Point this session at a different gateway. :gateway alone shows the current URL; :gateway clear reverts. | | :model <alias> | Switch the routing alias mid-session (auto, default, cheap, frontier, or a literal provider/model). :model alone shows the current value. | | :reset | Clear the screen and reset the turn counter to 1. Does not touch credentials. | | :smoke [N] [--model alias] [--limit N] | Run the smoketest harness without leaving the shell. The positional N and --limit are interchangeable; passing both with different values is an error. | | :testkey | Print a reminder of how to paste a real key. Does not itself change the active key. |

Dropping a file to batch

Paste or drag a .txt, .md, .prompt, or .prompts file path into the prompt -- as the only token on the line, resolving to a file that exists -- to dispatch every line in it as a batch, at concurrency 4, without leaving the shell:

iq[3]> ~/Desktop/support-queue.txt
dropped file support-queue.txt (23 prompts, model=auto, concurrency=4, 6 match canonical corpus)

  1/23  ✓  exp=-        obs=cheap    gpt-5-nano          cache-miss  0.61s  $0.00009
       complexity=simple (1.8/10)  focus=balanced  tokens=22→64  req=a1b2c3d4
       > "How do I reset my password?"
...

Aggregate
  Total prompts:        23
  ...

Your per-team rate limit depends on your plan: 240 requests per minute on Free, 600 on Team, and 6,000 or more on Enterprise. A separate, tighter burst ceiling applies per second, so a highly concurrent run can see a 429 while still comfortably under the per-minute number. Check your own limits on the rate-limits page in the dashboard.

The shell's own warning is a fixed heuristic rather than a reading of your plan: it warns above 60 prompts in a single drop, which is the most conservative tier's safe margin. Treat it as a nudge to check your limit, not as the limit itself. Re-dropping the same file replays every row from cache in a fraction of the original time. exp= shows an expected tier only for prompts that match the built-in smoketest corpus verbatim; anything else shows exp=- and only the observed tier.

Batch mode

$ iq batch prompts.txt --concurrency=4 --out=results.jsonl

Reading prompts.txt: 20 prompts loaded
Concurrency: 4

  1/20  ✓  simple    gpt-5-nano   cache-miss   0.84s   $0.00012
  2/20  ✓  medium    gpt-5-mini   cache-miss   1.92s   $0.00048
  3/20  ✓  simple    gpt-5-nano   cache-hit    0.04s   $0.00000
  ...
 20/20  ✓  complex   opus-4-8     cache-miss   3.41s   $0.00284

Aggregate
  Total prompts:        20
  Successes:            20
  Failures:              0
  Cache hit rate:       35%
  Tier mix:             simple 8 · medium 9 · complex 3
  Latency p50:        1.42s
  Latency p99:        4.18s
  Total spend:        $0.0341
  Avg cost / prompt:  $0.00171

  Wrote results.jsonl with 20 rows.

Flags:

  • --concurrency=N -- in-flight requests. Defaults to 4. Caps at 16 (the safe per-org ceiling); higher values clamp with a stderr warning.
  • --out=results.jsonl -- write one JSONL row per prompt with the full response body and the telemetry inlined under x_iq_routing.
  • --model, --timeout -- same semantics as iq prompt.

The exit code goes non-zero when at least one prompt fails (any 4xx or 5xx response, a timeout, or a transport error). Failed prompts do not abort the run; the aggregate reports successes and failures together. A 429 counts as a failure here too -- iq batch does not auto-retry a rate limit the way iq smoketest does, so lower --concurrency or split the file if a run is tripping your plan's limit.

With --out, a failed prompt still gets a row: response and x_iq_routing are null, and error holds the failure message.

Prompt-file format

One prompt per line. Blank lines and #-prefixed lines ignored.

# easy
What is the capital of France?
Translate "good morning" to Spanish.

# medium
Given this CSV row, output a JSON object with the columns as keys: 1,Alice,30

Prompt generator

iq generate-prompts --difficulty=mixed --count=20 --out=prompts.txt

Difficulty bands:

  • easy -- single-sentence factual lookups, single-step instructions. Target tier simple.
  • medium -- multi-step reasoning, structured output, short code snippets. Target tier medium.
  • hard -- long-context, ambiguous instructions, cross-domain synthesis. Target tier complex.
  • mixed -- 40% easy, 45% medium, 15% hard. A reasonable default for exercising all three tiers in one file.

Counts cap at 1000 per file. The generator does not call the gateway; it is a pure offline emitter.

Smoketest

iq smoketest
iq smoketest --limit 10 --model frontier

iq smoketest drives a canned corpus of 100-plus prompts through the gateway and checks that each one landed in its expected tier (cheap, default, or frontier). It is a routing-correctness check, not a synthetic dry run: every prompt is a real request against your provider keys, so a full run has a real, if small, cost. Prompts dispatch sequentially, so the full corpus takes several minutes.

$ iq smoketest --limit 3
iq smoketest  3 prompts  gateway=https://gateway.iq-routing.com  model=auto

  PASS sp-cheap-01     expected=cheap    model=gpt-5-nano                          84ms
       > What is the capital of France?
  PASS sp-default-01   expected=default  model=gpt-5.6-luna                      1204ms
       > Write a Python function that takes a list of integers and returns the second largest. Handle ties and empty lists.
  FAIL sp-frontier-01  expected=frontier model=gpt-5-mini                         932ms  (routed to 'default' tier; expected 'frontier' (model: gpt-5-mini))
       > Design the architecture for a multi-tenant SaaS analytics product handling 50 million events per day. Cover ingestion...

Routing distribution (expected tier x observed tier)
...

totals: 2 pass  1 fail  3 distinct chosen-model values  runtime 14.2s

Flags:

  • --limit N / -n N -- cap the run at the first N prompts (a head-of-list subset) instead of driving the full corpus.
  • --model / -m -- routing alias for the whole run: auto, cheap, default, or frontier, or a literal provider/model string. Same alias semantics as iq prompt --model; auto is used when this flag is omitted.
  • --timeout -- per-request timeout in seconds. Default 60.

The run fails (exit code 1) if any prompt's response is empty, missing telemetry, missing a chosen model, non-positive latency, missing cost on a non-cache-hit, or routed outside its expected tier -- or if the whole run produced fewer than four distinct chosen-model values. A 429 mid-run waits for the suggested retry window once and resumes rather than failing the prompt outright, since a full run can brush against your plan's rate limit. The X-RateLimit-Scope header on the 429 names which ceiling fired, and Retry-After gives the wait.

:smoke [N] [--model alias] [--limit N] runs the same harness inside the interactive shell without spawning a separate process; see Interactive shell, above.

Agent (dry-run planner)

iq agent "Find every TODO comment in the src/ directory and summarise them"
iq agent "Check whether api.example.com is up" --tools http_fetch

iq agent is a single-turn, dry-run planner. It sends your instruction plus a small catalog of advisory tool schemas (read_file, http_fetch) to the gateway and prints the model's proposed plan: either a numbered list of tool calls or a prose explanation.

The agent never executes anything. No file is read, no URL is fetched, no code runs. The tool schemas exist so the model can propose concrete, well-formed steps -- the operator reads the plan and runs whatever looks right by hand. Do not build automation on top of this command expecting the proposed steps to have actually happened; nothing downstream of the plan is real until you make it real.

$ iq agent "Find every TODO comment in the src/ directory and summarise them"
iq agent (planner / dry-run)  session=iqa_4f9c2e1a08b7c391  model=auto  tools=http_fetch,read_file
instruction: Find every TODO comment in the src/ directory and summarise them

Routing telemetry
  model=gpt-5-mini  cache-miss  0.81s  $0.00031  complexity=medium (5.1/10)  tokens=41→96

Proposed plan (advisory only -- nothing executed)
  1. read_file(path='src/index.ts')
  2. read_file(path='src/utils.ts')

Model commentary:
Run these two reads, then grep the results for "TODO" and summarise
each hit with its surrounding context.

Run any of these steps yourself if they look right; the agent never
executes tools.

Flags:

  • --model / -m -- routing alias or literal provider/model. Same semantics as iq prompt --model.
  • --tools -- comma-separated list of tool names the planner may propose. Built-ins: read_file, http_fetch. Defaults to both.
  • --theme -- light (default) or dark.

Every turn carries an X-IQ-Session-Id header (iqa_<16-hex-chars>), so the dashboard's Sessions view rolls the telemetry under one row even though the agent itself never loops.

Tokens, scopes, and revoke

The CLI authenticates with a personal access token issued from Settings → CLI access tokens. Default scopes:

  • read:routing -- view routing decisions.
  • write:prompt -- send prompts via the CLI.

Wider scopes (read:audit, write:settings) are opt-in at issue time. Tokens display the plaintext exactly once at issue. The dashboard shows only the prefix on subsequent reads.

iq logout removes the local credentials file. To revoke server-side, open Settings → CLI access tokens and click Revoke on the token row.

What bounds a token's spend

A personal access token is bounded at the organisation level, not at the credential level. Your org's daily and monthly budgets both apply to CLI traffic, and the daily budget is enforced before the call is made rather than after, so a token cannot spend past it. Per-token controls are a different story: unlike an org API key, a personal access token carries no per-credential spend ceiling and no per-credential rate limit of its own, so it inherits your team and plan defaults instead.

The practical effect is that a token is a narrower credential than an API key in what it can reach, and a blunter one in how finely you can cap it. If you need a hard per-credential ceiling, issue an org API key scoped to a team and set the ceiling there.

Bundle mode (Enterprise)

There is no self-hosted distribution on the self-serve plans. The standalone gateway bundle is the Enterprise on-prem option, arranged with us rather than downloaded, and it ships the CLI as part of the single-binary install. iq-gateway prompt "..." posts to the local gateway at http://localhost:8000; the same telemetry box renders.

iq-gateway serve            # boots the gateway + opens the dashboard
iq-gateway prompt "..."     # uses the local gateway

Some of the bundle's deeper analytics surfaces render a Limited mode banner: the on-prem install doesn't run the same aggregate reporting queries as the hosted product. Routing, caching, and per-request telemetry all run end to end in the bundle; only the aggregate analytics views are reduced.

Troubleshooting

  • Token in {path} is not a recognised IQ credential -- the saved value carries no known prefix. Re-run iq login and paste an org API key (gw_live_...) or a CLI access token (iqp_...).
  • Credentials file is group- or world-readable -- POSIX safety check. Run chmod 0600 ~/.config/iq-routing/credentials.toml.
  • Gateway returned 401 -- the token expired or was revoked. Run iq login to reissue.
  • Telemetry box shows telemetry unavailable -- the gateway you are pointed at is older than the telemetry header, or telemetry has been turned off for that deployment. Update to a current gateway version, or re-enable the telemetry header on the deployment you are calling.

See also

  • /docs/quickstart -- SDK integration via OpenAI/Anthropic clients.
  • /docs/api-reference -- REST surface details.
  • /integrate -- drop-in code snippets and a Use the IQ CLI panel.