Skip to content
IQ Routing

Prompt token optimizer

The token optimizer is a payload-size pass on the user-side messages and the system prompt before the request reaches the upstream provider. A 30 percent compression on a Sonnet call is roughly 30 percent off the input token cost on that request. The optimizer is a FinOps lever, not a quality lever; it does not rewrite prompts to be "better."

If anything goes wrong while optimizing, IQ sends your original prompt unchanged, so the optimizer never blocks or delays a request.

Three modes

| Mode | What it does | Latency budget | When to use | |---|---|---|---| | off | Pass-through. No compression. | 0 ms | Default. Use when payload size is not a cost driver. | | safe | Deterministic, lossless cleanup of redundant formatting in conversational text. | Negligible -- no model call, so it never adds meaningful latency. | Default-on recommendation. Lossless for the model; observable token reduction on chat-heavy prompts. | | aggressive | Aggressive mode removes more low-value text from the prompt for a larger token reduction, with a small risk to answer quality. | 250 ms hard cap; fails open above. | Cost-driven workloads. Gated by an explicit opt-in so the trade-off is a conscious choice, not a default. |

What the optimizer never touches

  • Tool-call payloads. Function arguments and structured outputs pass through unchanged. The compression scope is conversational text, not machine-readable JSON.
  • Assistant messages. The optimizer compresses what you send; what the model said back stays as-is for the conversation history.
  • Code fences. Triple-backtick blocks pass through whitespace- preserving even in safe mode. Code semantics depend on indentation.

Picking a mode

Settings → Token optimizer → mode picker. The aggressive option is a deliberate control: an explicit opt-in checkbox turns it on at the org's threshold, so choosing more compression for cost-driven workloads is a conscious choice rather than a default.

One composition to know about: an org whose routing focus mode is set to quality runs safe even when the optimizer dial says aggressive. An operator who asked for quality on the routing side does not silently get lossy compression on the prompt side.

The dashboard surface lives at /settings#token-optimizer.

The pane also exposes:

  • Threshold (aggressive only). The retain-or-drop threshold the aggressive pass uses. A lower threshold means more aggressive compression and a higher quality risk.
  • Compression stats. A 30-day rollup of tokens saved, cost saved, average compression ratio, and the number of requests the optimizer ran on. Tokens saved is the number the ROI case rests on.

Safe mode

Safe mode is a deterministic, lossless cleanup of redundant formatting in conversational text. It tidies the kind of structure that costs tokens without carrying meaning for the model, and it preserves the content of the prompt verbatim. Code fences pass through whitespace-preserving so indentation-sensitive content is never touched. The result is a smaller payload with no change to what the model reads as substance.

Aggressive mode

Aggressive mode removes more low-value text from the prompt for a larger token reduction, with a small risk to answer quality. If it can't finish quickly or isn't available, IQ uses safe mode's cleanup or sends the prompt unchanged. The standalone bundle supports safe mode only.

How much a prompt can shrink

Each mode limits how much it will shrink a prompt. If a result would go past that limit, IQ sends your original prompt and records a compression_threshold_breached event, which you can subscribe to with webhooks. Aggressive mode allows more reduction than safe mode.

Rollback

Rollback is a mode change, not a deploy. If safe mode misbehaves on your traffic, set the org's optimizer mode to off in the dashboard and the pass-through path restores the prior behaviour on the next request. If aggressive mode misbehaves, drop the org back to safe from the same picker; the aggressive pass stops running immediately and nothing else about the request path changes.

Bundling for offline use (Enterprise)

The standalone gateway bundle is the Enterprise on-prem option, arranged with us rather than downloaded; there is no self-hosted distribution on the self-serve plans.

The bundle does not ship the aggressive-mode weights and does not load them at runtime, on any host, regardless of network access. An org running the bundle that is set to aggressive runs safe instead -- the same gateway-wide fallback described in Aggressive mode, above. safe mode is unaffected: it is pure Python with no model to load, so it runs identically in the bundle and on the hosted gateway.

See also

  • /docs/billing -- token classes and the four-way cost split the optimizer's savings show up against.
  • /docs/caching -- the optimizer runs before cache lookup, so a compressed prompt has its own cache key, and a gateway cache hit is billed at zero whether or not the optimizer ran.
  • /docs/webhooks -- subscribe to compression_threshold_breached to get notified the moment a floor breach ships an uncompressed prompt, instead of finding out from the 30-day compression-stats rollup.
  • /docs/quickstart -- the SDK setup does not need to change for the optimizer to apply.