Skip to content
IQ Routing

LangChain · Python

LangChain's ChatOpenAI and ChatAnthropic wrappers accept the same base_url argument as the underlying SDK. Swap the base URL on the model constructor and every chain that uses the LLM routes through the gateway.

Drop-in switch

from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="auto",
    base_url="https://gateway.iq-routing.com/v1",
    api_key="gw_live_xxxxxxxx",
)

See Authentication for how to generate a gw_live_... key and where to store it.

langchain-anthropic follows the same pattern with base_url pointing at the gateway origin (no /v1 suffix; LangChain's wrapper appends the path itself).

Verify it routes

import os
from langchain_openai import ChatOpenAI
from langchain_core.messages import HumanMessage

llm = ChatOpenAI(
    model="auto",
    base_url="https://gateway.iq-routing.com/v1",
    api_key=os.environ["IQ_API_KEY"],
    timeout=60,
    max_retries=0,
)

result = llm.invoke([HumanMessage(content="Summarise: routing-aware gateway.")])
print("Content:", result.content)
print("Model picked:", result.response_metadata.get("model_name"))
print("Token usage:", result.response_metadata.get("token_usage"))

Drop into a fresh venv with pip install langchain langchain-openai. Set IQ_API_KEY in your environment, run. The response_metadata["model_name"] field carries the chosen upstream model; the token_usage dict matches the standard OpenAI completion tokens shape.

The Anthropic equivalent:

from langchain_anthropic import ChatAnthropic

llm = ChatAnthropic(
    model="claude-haiku-4-5-20251001",
    base_url="https://gateway.iq-routing.com",
    api_key=os.environ["IQ_API_KEY"],
    timeout=60,
    max_retries=0,
)

result = llm.invoke([HumanMessage(content="Say hello.")])

Using capability aliases

The cap:<name> syntax in the model field tells the gateway to pick a concrete model at routing time based on the capability's stable intent rather than a pinned model id. Your org's per-capability overrides and current provider health can shift which concrete model a capability resolves to on any given call; the capability's intent stays stable (reason-heavy always favors a stronger reasoning model). Five default capabilities resolve to a concrete model this way: reason-heavy, tool-call-strict, long-context-128k, vision, and json-mode. A sixth, cheap-fast, is accepted too but currently routes like auto -- see capability aliases.

from langchain_anthropic import ChatAnthropic

chat = ChatAnthropic(
    model="cap:reason-heavy",
    anthropic_api_url="https://gateway.iq-routing.com",
    api_key=os.environ["IQ_API_KEY"],
)

result = chat.invoke([HumanMessage(content="Plan the migration in three phases.")])

The resolved provider plus model surface in the x-iq-routing response header, but LangChain does not expose that raw header (see the gotchas below). The chain author can still read the resolved model through response_metadata["model_name"], which every .invoke() result populates. See the capability aliases docs for the full list of default mappings and how to override them. You can override the default capability mapping for your org in the dashboard at /settings/capabilities.

Common gotchas

A missing or invalid API key returns 401 before the request reaches routing -- no provider call is made and nothing is billed. LangChain surfaces this as the underlying SDK's auth exception; see the API reference for the full set of status codes and error body shapes.

LangChain wraps the underlying SDK but does not expose the raw response headers by default, so the X-Request-Id header itself is not reachable through response_metadata or a callback. The gateway does embed its request id inside the completion, though: the chat-completion id field comes back as chatcmpl-<request-id>, one chatcmpl- prefix away from the value /requests/<request-id> expects. Whether that raw id survives into response_metadata or llm_output depends on your langchain-openai version, so treat it as a best-effort capture rather than a guaranteed one:

def request_id_from_response(result) -> str | None:
    raw_id = (result.llm_output or {}).get("id") or getattr(result, "id", None)
    if raw_id and raw_id.startswith("chatcmpl-"):
        return raw_id.removeprefix("chatcmpl-")
    return None

If correlation matters for every call, call the underlying SDK directly instead of going through LangChain -- see the OpenAI Python recipe for the with_raw_response pattern that reads X-Request-Id off the real HTTP response.

LangChain's default retry-on-429 is aggressive (up to 6 retries with exponential backoff). The gateway's 429 carries Retry-After; LangChain's retries do not respect the header. Set max_retries=0 on the model constructor and rely on your own queue.

LangSmith tracing (LANGCHAIN_TRACING_V2=true) captures full prompt content. If you route through the gateway, every prompt traces to LangSmith's servers in addition to the gateway's own routing records. Pin your tracing scope; the gateway alone has the routing decision, the costs, and the cache hits.

For ChatAnthropic, the gateway forwards the anthropic-beta header verbatim rather than evaluating it, so a beta your chain sets reaches Anthropic unchanged. If Anthropic rejects the beta the gateway surfaces that 400 with a request_id; no silent fall-through.

Try this in the playground

Run this snippet against your gateway key →