OpenAI · Node
The OpenAI Node/TypeScript SDK accepts baseURL and apiKey on the client
constructor. Swap the base URL and your existing call sites route through the
gateway. This recipe ships a runnable TypeScript file that you can drop into
a fresh pnpm init project with openai installed.
Drop-in switch
const client = new OpenAI({
baseURL: "https://gateway.iq-routing.com/v1",
apiKey: process.env.IQ_API_KEY,
});
IQ_API_KEY holds your gateway key (gw_live_xxxxxxxx). See
Authentication for how to generate one and where to store it.
Everything below this line (client.chat.completions.create,
client.models.list) routes through the gateway. The gateway does not
implement /v1/embeddings, so leave embedding calls pointed at your
provider.
Verify it routes
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://gateway.iq-routing.com/v1",
apiKey: process.env.IQ_API_KEY!,
timeout: 60_000,
maxRetries: 0,
});
const response = await client.chat.completions
.create({
model: "auto",
messages: [{ role: "user", content: "Summarise: routing-aware gateway." }],
max_tokens: 64,
})
.withResponse();
console.log("X-Request-Id:", response.response.headers.get("x-request-id"));
const telemetry = JSON.parse(response.response.headers.get("x-iq-routing") ?? "{}");
console.log("Model picked:", telemetry.chosen_model);
console.log("Provider:", telemetry.chosen_provider);
console.log("Cache hit:", telemetry.cache_hit);
console.log(response.data.choices[0].message.content);
Run with IQ_API_KEY=... npx tsx verify.ts. The gateway emits two
response headers per request: X-Request-Id (the correlation handle) and
x-iq-routing (a JSON-encoded telemetry blob with the chosen model,
provider, complexity tier, cache-hit flag, latencies, and cost). The
body is the standard OpenAI chat completion shape.
Using capability aliases
The cap:<name> syntax in the model field tells the gateway to pick a
concrete model at routing time based on the capability's stable
intent rather than a pinned model id. Your org's per-capability
overrides and current provider health can shift which concrete model a
capability resolves to on any given call; the capability's intent
stays stable (reason-heavy always favors a stronger reasoning model).
Five default capabilities resolve to a concrete model this way:
reason-heavy, tool-call-strict, long-context-128k, vision, and
json-mode. A sixth, cheap-fast, is accepted too but currently
routes like auto -- see capability aliases.
import OpenAI from "openai";
const openai = new OpenAI({
baseURL: "https://gateway.iq-routing.com/v1",
apiKey: process.env.IQ_API_KEY!,
});
const completion = await openai.chat.completions.create({
model: "cap:reason-heavy",
messages: [{ role: "user", content: "Plan the migration in three phases." }],
});
The resolved provider plus model surface in the x-iq-routing response
header (a JSON-encoded payload with chosen_provider and chosen_model
fields, among others), so the SDK consumer can inspect which concrete
model handled the request via .withResponse(). Decode the header with
JSON.parse(response.response.headers.get("x-iq-routing")). See the
capability aliases docs for the full list of
default mappings and how to override them. You can override the
default capability mapping for your org in the dashboard at
/settings/capabilities.
Common gotchas
A missing or invalid Authorization header returns 401 before the
request reaches routing -- no provider call is made and nothing is
billed. See the API reference for the full set
of status codes and error body shapes.
Node's default fetch timeout depends on the runtime (Node 18+ uses
undici, which has a default of 5 minutes for headers). The gateway's own
worst case is longer still: up to 600 seconds per provider attempt and 1800
seconds (30 minutes) across a full fallback chain before it gives up with a
504. Most requests finish in a few seconds; pass an explicit timeout on
the client if your product needs to fail faster than either ceiling.
When every upstream provider in the fallback chain is exhausted before
that ceiling is hit, the gateway returns 503 with an overloaded_error
type in the response body, and no Retry-After header -- back off on
your own schedule for that case.
The SDK's maxRetries defaults to 2. The gateway's 429 carries
a Retry-After header set to the exact cooldown; SDK-side retries on
top of that just burn calls inside the same cooldown window. Set
maxRetries: 0 and let your own queue handle backoff.
Streaming uses the standard stream: true flag. The gateway forwards the
chunked SSE; the X-Request-Id arrives on the response headers before the
first chunk, so you can log the correlation handle even on a streamed
response.
If you use the SDK in a Next.js route handler, set runtime: "nodejs" on
the route. Edge runtime works for the request itself but loses the
streaming SSE shape on some Vercel regions; the Node runtime is the safe
default for the gateway's response shape.