LangChain · Python
LangChain's ChatOpenAI and ChatAnthropic wrappers accept the same
base_url argument as the underlying SDK. Swap the base URL on the model
constructor and every chain that uses the LLM routes through the gateway.
Drop-in switch
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="auto",
base_url="https://gateway.iq-routing.com/v1",
api_key="gw_live_xxxxxxxx",
)
See Authentication for how to generate a gw_live_...
key and where to store it.
langchain-anthropic follows the same pattern with base_url pointing at
the gateway origin (no /v1 suffix; LangChain's wrapper appends the
path itself).
Verify it routes
import os
from langchain_openai import ChatOpenAI
from langchain_core.messages import HumanMessage
llm = ChatOpenAI(
model="auto",
base_url="https://gateway.iq-routing.com/v1",
api_key=os.environ["IQ_API_KEY"],
timeout=60,
max_retries=0,
)
result = llm.invoke([HumanMessage(content="Summarise: routing-aware gateway.")])
print("Content:", result.content)
print("Model picked:", result.response_metadata.get("model_name"))
print("Token usage:", result.response_metadata.get("token_usage"))
Drop into a fresh venv with pip install langchain langchain-openai. Set
IQ_API_KEY in your environment, run. The
response_metadata["model_name"] field carries the chosen upstream model;
the token_usage dict matches the standard OpenAI completion tokens
shape.
The Anthropic equivalent:
from langchain_anthropic import ChatAnthropic
llm = ChatAnthropic(
model="claude-haiku-4-5-20251001",
base_url="https://gateway.iq-routing.com",
api_key=os.environ["IQ_API_KEY"],
timeout=60,
max_retries=0,
)
result = llm.invoke([HumanMessage(content="Say hello.")])
Using capability aliases
The cap:<name> syntax in the model field tells the gateway to pick a
concrete model at routing time based on the capability's stable
intent rather than a pinned model id. Your org's per-capability
overrides and current provider health can shift which concrete model a
capability resolves to on any given call; the capability's intent
stays stable (reason-heavy always favors a stronger reasoning model).
Five default capabilities resolve to a concrete model this way:
reason-heavy, tool-call-strict, long-context-128k, vision, and
json-mode. A sixth, cheap-fast, is accepted too but currently
routes like auto -- see capability aliases.
from langchain_anthropic import ChatAnthropic
chat = ChatAnthropic(
model="cap:reason-heavy",
anthropic_api_url="https://gateway.iq-routing.com",
api_key=os.environ["IQ_API_KEY"],
)
result = chat.invoke([HumanMessage(content="Plan the migration in three phases.")])
The resolved provider plus model surface in the x-iq-routing
response header, but LangChain does not expose that raw header (see
the gotchas below). The chain author can still read the resolved
model through response_metadata["model_name"], which every
.invoke() result populates. See the capability aliases
docs for the full list of default mappings and
how to override them. You can override the default capability mapping
for your org in the dashboard at /settings/capabilities.
Common gotchas
A missing or invalid API key returns 401 before the request reaches
routing -- no provider call is made and nothing is billed. LangChain
surfaces this as the underlying SDK's auth exception; see the
API reference for the full set of status codes
and error body shapes.
LangChain wraps the underlying SDK but does not expose the raw
response headers by default, so the X-Request-Id header itself is not
reachable through response_metadata or a callback. The gateway does embed
its request id inside the completion, though: the chat-completion id
field comes back as chatcmpl-<request-id>, one chatcmpl- prefix away
from the value /requests/<request-id> expects. Whether that raw id
survives into response_metadata or llm_output depends on your
langchain-openai version, so treat it as a best-effort capture rather
than a guaranteed one:
def request_id_from_response(result) -> str | None:
raw_id = (result.llm_output or {}).get("id") or getattr(result, "id", None)
if raw_id and raw_id.startswith("chatcmpl-"):
return raw_id.removeprefix("chatcmpl-")
return None
If correlation matters for every call, call the underlying SDK directly
instead of going through LangChain -- see the
OpenAI Python recipe for the
with_raw_response pattern that reads X-Request-Id off the real HTTP
response.
LangChain's default retry-on-429 is aggressive (up to 6 retries with
exponential backoff). The gateway's 429 carries Retry-After;
LangChain's retries do not respect the header. Set
max_retries=0 on the model constructor and rely on your own queue.
LangSmith tracing (LANGCHAIN_TRACING_V2=true) captures full prompt
content. If you route through the gateway, every prompt traces to
LangSmith's servers in addition to the gateway's own routing
records. Pin your tracing scope; the gateway alone has the routing decision,
the costs, and the cache hits.
For ChatAnthropic, the gateway forwards the anthropic-beta header
verbatim rather than evaluating it, so a beta your chain sets reaches
Anthropic unchanged. If Anthropic rejects the beta the gateway surfaces
that 400 with a request_id; no silent fall-through.