Savings methodology
Every savings figure on this site comes from one of two places, and this page names which. The first is a worked five-step agent loop where each per-step price is printed on the home page, so the total is arithmetic you can check rather than a claim you have to trust. The second is measured traffic through the production gateway, priced from the gateway’s own billing table, and that is where the 40 to 80 percent band in the headline comes from. Neither is a guarantee for your workload, because a saving is always a saving against some baseline, and the baseline is the part most vendors leave out.
A saving is only a saving against a baseline
The number that matters is not how much IQ Routing saves, it is how much it saves against the thing you would otherwise have done. Across 1,071 requests of organic traffic through the production gateway, priced from the gateway’s own billing table with no synthetic load, routed spend came in 66 percent under pinning every request to a top-tier reasoning model, and 43 percent under pinning every request to a mid-tier Sonnet-class model. Against pinning every request to gpt-5, the same traffic came in 27 percent more expensive.
That last figure is the one worth reading twice, and we publish it on purpose. Routing is not free, and a router that always wins is a router choosing its baseline after the fact. IQ Routing pays for itself when your default model is more expensive than most of your steps need, which is the ordinary shape of an agent loop, and it does not pay for itself when your default is already sized close to the work. Overall, the honest version of the pitch is that we will tell you which of those two you are, and the dashboard reports your routed cost against the baseline you choose rather than the one that flatters us.
The 43 percent figure sets the floor of the 40 to 80 percent band on the home page, rounded down to the nearest ten. The top comes from a second, larger measurement, on the same gateway billing table but a different, named baseline: $5 per million input tokens and $27.50 per million output tokens, the midpoint of Claude Opus 4.8 ($5/$25) and OpenAI’s gpt-5.6-sol ($5/$30), the savings dashboard’s blend throughout that window, in place since July 10, 2026. Over the 30 days ending August 19, 2026, our account routed 6,530 requests and came in 80.6 percent under that blended baseline. Eighty is 80.6 rounded down to the nearest ten, the same rule the floor uses: neither end of the band claims more than was measured. The savings dashboard has since moved its own baseline to a blend of Claude Opus 4.8 and GPT-6 Astra, which reads higher on the same traffic; this band keeps the older, lower reading.
Both readings share the same limitation, stated plainly. Each one re-prices the tokens a request actually used at its own baseline price; neither re-runs the request against that baseline model, which could tokenize, retry, or answer at a different length than the model that actually served it. It is a real bill compared against published list prices, not a live rehearsal. Which end of the band you land on also depends on how expensive your current default is relative to the work your steps actually do, and the gpt-5 result above is the reminder that a default already sized close to the work does not land inside the band at all. Overall, the band is a rounded envelope of two measured snapshots rather than a target we picked first and then went looking for evidence to support.
One caveat on both samples above, stated plainly. The 1,071-request and 6,530-request figures are our own accounts rather than customer traffic; we are signing our first design partners now. The two samples were read at different times and may share some of the same requests, so read the band as two snapshots rather than one continuous curve. Both are real measurements of a real workload rather than a projection, but it is our workload, and we would rather say so than dress it up as somebody else’s.
The 57 percent before and after example
The before and after example on the home page is a single five-step agent loop: planning, retrieval, a tool call, synthesis, and verification. In the Before trace every step runs on the same top-tier reasoning model at high thinking effort, and the five step prices of $0.21, $0.19, $0.17, $0.39, and $0.35 total $1.31 for the loop. In the After trace IQ Routing resolves each step to the cheapest model that clears the quality bar for that step, so synthesis stays on the same heavy reasoning model at $0.39, planning moves to a cheaper reasoning-heavy model at $0.11, and retrieval, the tool call, and verification drop to cheaper, faster variants at $0.02, $0.03, and $0.01. Those five total $0.56, so the loop gives back $0.75 of $1.31, which is 57 percent.
This is a worked example rather than a customer result, and it is one representative loop chosen to show the mechanism. The baseline is named in the trace itself: every Before step runs on the same top-tier reasoning model, which is the configuration most agent frameworks ship with. The per-step prices reflect a specific mix of providers and step types, so a loop with a different shape, a different provider mix, or a higher share of reasoning-heavy steps will land at a different number.
Two different things get called a cache hit
Conflating them is how a cache rate stops meaning anything, so here is the split. The first is IQ Routing’s own response cache, where the gateway recognises a repeat request and returns the stored answer without calling a provider at all, so the request costs nothing and comes back immediately. Across 1,137 organic requests, a larger sample of the same production traffic behind the routing figures above, 163 were served that way, which is 14.3 percent. The second is provider prompt caching, where the provider still runs the model but bills repeated input tokens at a tenth of the normal input rate. Measured as cache-read input tokens over total input tokens presented, that comes to 49.1 percent, and it is the larger of the two effects.
That 49.1 percent carries a date, and the date is part of the number. Provider cache accounting only shipped on July 20, 2026, so the measurement runs from July 21 onward, and ten of the fourteen active days in the sample record zero cache reads because nothing was recording them rather than because nothing was cached. What survives that cut is four days, 200 billed requests, and a single organisation. It is a real measurement rather than a projection, and it is a narrow one, so read it as an early rate rather than a settled one.
Combined against a no-cache counterfactual, the two effects avoided 17.4 percent of what the same traffic would otherwise have cost over the whole period from June 24 to July 28. One gap in that accounting is ours to name rather than yours to find: cache creation tokens are not yet recorded, and writing to a provider cache costs slightly more than an ordinary input token, so both percentages above are marginally flattering and will settle a little lower once the write side is instrumented. Overall, both numbers belong on this page, the small one beside the large one, because a cache rate quoted without naming which cache produced it tells a buyer nothing, and we would rather publish the 14.3 percent than let the 49.1 percent do work it has not earned.
Why results vary by workload
Three factors drive most of the variance. The first is model mix: the more of your traffic that can run on a cheaper model without losing quality, the larger the cut. The second is cache hit rate: repeated or near-repeated prompts served from the semantic cache cost nothing on the provider side, so workloads with high reuse save more, while first-of-a-kind prompts save less. The third is step composition: agent loops with many light steps (retrieval, tool calls, short formatting) leave more headroom to route down than loops dominated by hard reasoning. Your own savings depend on how these three land for your traffic.
How to read these numbers
The figures on this site are a mix of worked examples and measurements of our own production traffic, and each section above says which one it is rather than blending them into a single number. The dashboard reports your own routed cost against a single-model baseline, so you see the exact figure for your traffic. The way to size the savings for your workload is to run it through IQ Routing and read the measured result, which scales with your model mix, cache hit rate, and step composition.