Skip to content
IQ Routing

OpenAI · Node

The OpenAI Node/TypeScript SDK accepts baseURL and apiKey on the client constructor. Swap the base URL and your existing call sites route through the gateway. This recipe ships a runnable TypeScript file that you can drop into a fresh pnpm init project with openai installed.

Drop-in switch

const client = new OpenAI({
  baseURL: "https://gateway.iq-routing.com/v1",
  apiKey: process.env.IQ_API_KEY,
});

IQ_API_KEY holds your gateway key (gw_live_xxxxxxxx). See Authentication for how to generate one and where to store it.

Everything below this line (client.chat.completions.create, client.models.list) routes through the gateway. The gateway does not implement /v1/embeddings, so leave embedding calls pointed at your provider.

Verify it routes

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://gateway.iq-routing.com/v1",
  apiKey: process.env.IQ_API_KEY!,
  timeout: 60_000,
  maxRetries: 0,
});

const response = await client.chat.completions
  .create({
    model: "auto",
    messages: [{ role: "user", content: "Summarise: routing-aware gateway." }],
    max_tokens: 64,
  })
  .withResponse();

console.log("X-Request-Id:", response.response.headers.get("x-request-id"));
const telemetry = JSON.parse(response.response.headers.get("x-iq-routing") ?? "{}");
console.log("Model picked:", telemetry.chosen_model);
console.log("Provider:", telemetry.chosen_provider);
console.log("Cache hit:", telemetry.cache_hit);
console.log(response.data.choices[0].message.content);

Run with IQ_API_KEY=... npx tsx verify.ts. The gateway emits two response headers per request: X-Request-Id (the correlation handle) and x-iq-routing (a JSON-encoded telemetry blob with the chosen model, provider, complexity tier, cache-hit flag, latencies, and cost). The body is the standard OpenAI chat completion shape.

Using capability aliases

The cap:<name> syntax in the model field tells the gateway to pick a concrete model at routing time based on the capability's stable intent rather than a pinned model id. Your org's per-capability overrides and current provider health can shift which concrete model a capability resolves to on any given call; the capability's intent stays stable (reason-heavy always favors a stronger reasoning model). Five default capabilities resolve to a concrete model this way: reason-heavy, tool-call-strict, long-context-128k, vision, and json-mode. A sixth, cheap-fast, is accepted too but currently routes like auto -- see capability aliases.

import OpenAI from "openai";

const openai = new OpenAI({
  baseURL: "https://gateway.iq-routing.com/v1",
  apiKey: process.env.IQ_API_KEY!,
});

const completion = await openai.chat.completions.create({
  model: "cap:reason-heavy",
  messages: [{ role: "user", content: "Plan the migration in three phases." }],
});

The resolved provider plus model surface in the x-iq-routing response header (a JSON-encoded payload with chosen_provider and chosen_model fields, among others), so the SDK consumer can inspect which concrete model handled the request via .withResponse(). Decode the header with JSON.parse(response.response.headers.get("x-iq-routing")). See the capability aliases docs for the full list of default mappings and how to override them. You can override the default capability mapping for your org in the dashboard at /settings/capabilities.

Common gotchas

A missing or invalid Authorization header returns 401 before the request reaches routing -- no provider call is made and nothing is billed. See the API reference for the full set of status codes and error body shapes.

Node's default fetch timeout depends on the runtime (Node 18+ uses undici, which has a default of 5 minutes for headers). The gateway's own worst case is longer still: up to 600 seconds per provider attempt and 1800 seconds (30 minutes) across a full fallback chain before it gives up with a 504. Most requests finish in a few seconds; pass an explicit timeout on the client if your product needs to fail faster than either ceiling.

When every upstream provider in the fallback chain is exhausted before that ceiling is hit, the gateway returns 503 with an overloaded_error type in the response body, and no Retry-After header -- back off on your own schedule for that case.

The SDK's maxRetries defaults to 2. The gateway's 429 carries a Retry-After header set to the exact cooldown; SDK-side retries on top of that just burn calls inside the same cooldown window. Set maxRetries: 0 and let your own queue handle backoff.

Streaming uses the standard stream: true flag. The gateway forwards the chunked SSE; the X-Request-Id arrives on the response headers before the first chunk, so you can log the correlation handle even on a streamed response.

If you use the SDK in a Next.js route handler, set runtime: "nodejs" on the route. Edge runtime works for the request itself but loses the streaming SSE shape on some Vercel regions; the Node runtime is the safe default for the gateway's response shape.

Try this in the playground

Run this snippet against your gateway key →