Route every call to the cheapest model that holds quality.
IQ Routing drops in front of chatbots, RAG pipelines, agent loops, and finance workloads. On our own traffic it cut spend 40 to 80 percent; what you save depends on your default model and workload.
Card required. No charge on Free.
Drop in front of the OpenAI or Anthropic SDK you already use
Any OpenAI- or Anthropic-shaped endpoint drops in behind one base URL. See the full lineup of supported providers on our integrations page.
The problem
Every call hits your top model. The bill keeps climbing and nobody can say why.
Pin one frontier model on classification, retrieval, tool calls, and cleanup alike and you pay reasoning rates for work a cheap model would nail. Then the invoice lands as one line, shipping outruns anything finance can categorise, the cost report is a guess, and the audit trail is a Slack thread.
One worked example, before and after
The same five-step LangChain loop, routed two ways.
The Before trace runs every step on the same top-tier reasoning model, which totals $1.84. The After trace sends the same loop through IQ Routing, which totals $0.77, or 58 percent less for the same five steps. Every per-step price is on the page, so the arithmetic is yours to check rather than ours to assert, and nothing in the agent code changes.
Before · vanilla LangChain
one provider, one budgetEvery step on claude-opus-4-8 at thinking=high
After · IQ Routing
routed by IQ, one sessionEach step runs on a model that fits it
The fix
One gateway. Six superpowers.
Each one shows up on the dashboard the moment your first request lands.
Works inside the apps that lock models.
Some tools only allow one model family. Point Claude Code, Cursor, or your own app at IQ and requests stay inside the family the tool requires; subagents use another provider only if you turn that on. IQ holds the conversation too, so switching models mid-thread stays safe.
Read the docs →
Repeats resolve before they bill.
When a request matches one IQ has already answered, the answer comes straight back without another model call. IQ catches exact repeats and ones that mean the same thing. Each organization's cache is kept separate from every other organization's, and with zero data retention on, IQ doesn't cache your traffic at all.
Read the docs →
The router is the hard part.
Routing well is the whole product. Choose how IQ routes (Auto, Optimal, Cheap, Frontier, Dynamic, or your own table), and every request goes to a model from the providers you've connected, with automatic fallback if one has a problem. You see which model served each request and what it cost.
Read the docs →
Every dollar has an owner.
Per-team budgets, per-key limits, and a four-way input, output, cached, and thinking cost split, written to the ledger and enforced before the provider call.
Read the docs →
A session is one accountable unit.
Every multi-step agent loop tracks per-turn cost, latency, tokens, and step class, and a governor caps how much thinking the session can spend.
Read the docs →
Each step gets the right model.
An agent runs many steps, and most of them don't need your most expensive model. IQ sizes up each one and routes it to the cheapest model that can still do it well, so you get frontier quality where it counts and pay a fraction everywhere else.
Read the docs →
What the dashboard looks like
Spend, budgets, and per-team usage on one page.
Illustrative dashboard, sample data.
Pricing
Priced so it pays for itself.
Start free in thirty seconds. Scale into Team and Enterprise when the bill makes the case.
Free
$0 forever. Card on file, no trial, no charge to start.
One URL for every provider. Bring your own provider keys (BYOK).
- ✓OpenAI, Anthropic, and Google, with one URL for all three
- ✓Bring your own keys (BYOK)
- ✓Semantic cache
- ✓Email support
Team
RecommendedBilled monthly, no annual plan
Everything in Free plus alerting, member roles, and custom band maps.
- ✓Everything in Free
- ✓Per-team budgets
- ✓Audit log
- ✓Per-team alerting
- ✓Owner, admin and member roles
- ✓Custom band maps (auto, cheap, frontier)
- ✓Priority email support
Enterprise
Custom annual contract
On-prem deployment, SOC2 evidence pack on request, dedicated support.
- ✓Everything in Team
- ✓On-prem or VPC deployment
- ✓SOC2 evidence pack on request
- ✓ERP integrations (on the roadmap)
- ✓Dedicated support engineer
Three lines today. A clean bill tomorrow.
Route real traffic in under five minutes.
Card required. No charge on Free.