Skip to content
IQ Routing

Security

Last reviewed: August 12, 2026.

This page covers the per-org security controls surfaced in the Settings → Security pane, and it carries the authoritative statement of the Zero Data Retention boundary. Live in production today: the Zero Data Retention setting, the per-org PII redactor that strips sensitive content out of the prompt before it leaves the gateway, the anomaly rate-limit that caps how loudly the detector can fire, and the access scoping that gates every admin surface. One control is on the roadmap and documented here as designed, not yet live: the multi-region data-residency control that would pin the org's region to matching providers.

The audience is an org admin tuning security plus data-residency posture. Each section names the control the admin flips, the gateway behaviour the flip changes, and the roadmap work that will close the section out.

Zero data retention

Status: Live today. Per-org Zero Data Retention is in force in production and is self-serve on every plan, including Free.

Zero Data Retention (ZDR) is an org-level setting, flipped by an owner or an admin from the Settings pane, and it takes hold on the very next request. When it goes on, the gateway stamps an enablement timestamp server-side, so you have a fixed, non-client-supplied date from which the boundary below applies. This section is the authoritative statement of that boundary, written so you can hand it to your own security reviewer without paraphrasing it first.

The retention boundary

With ZDR on, payload logging and telemetry content retention are disabled and locked for your org. Request and response bodies are never written to durable storage: not to the requests browser, not to the audit trail, not by any operational record the gateway keeps. The lock is a real lock rather than a default, so an attempt to arm payload capture while ZDR is on is rejected outright with a 400 instead of being quietly honoured, and an already-open capture window is closed by the same flip rather than merely being blocked from renewing. Any free-text reasoning generated internally while routing your request, which is the one remaining place a fragment of your content could have surfaced inside an operational record, is withheld from telemetry as well.

Conversation state on the stateful API still exists, because that API cannot function without it, but it exists only to operate that API and it is bounded by a time-to-live. An idle conversation expires automatically, 24 hours by default, and an org-level erasure sweeps your conversations immediately rather than waiting for the TTL to run out. Nothing outside that mechanism keeps a copy.

With ZDR on, IQ Routing does not cache your organization's requests or responses: the gateway neither reads from nor writes to its response caches for your organization, including the semantic-similarity layer and its numeric fingerprint (an embedding) of each prompt. Entries cached before you turned ZDR on are no longer served; they expire on their normal schedule, within seven days, or an owner or admin can remove them at once with Clear my org's cache. See /docs/caching for how the cache layers work.

ZDR still allows operational metadata, including timing, token counts, cost and model choice, for billing and the requests browser. When loop detection is enabled for your organization, we also keep secret-key fingerprints of message and tool activity, activity counts, loop flags and estimated wasted cost to detect agents stuck in repeated loops. This applies with ZDR on too. Loop detection does not store message text, tool names, tool arguments or tool results. Its temporary activity summary expires 30 minutes after its last update. Loop metadata attached to request records has no automatic expiry. Diagnostic logs also retain loop metadata and organization and request identifiers.

What flipping it back off does

Turning ZDR off removes the lock and lets payload capture be enabled again on request, the same as before. It does not restore, backfill, or reconstruct anything that was never written while the lock was on. In the same way, turning ZDR on is forward-only: it stops new content from being recorded, and it does not reach back and delete payloads captured before the enablement timestamp. If you need those removed, use the erasure path rather than assuming the flip did it for you.

What retained payloads are used for

Payload logging exists to give your team a requests browser, not to feed a training pipeline. IQ Routing keeps an internal, manually-run export an engineer can invoke to build training data for the systems that select a model for each request, but only for an organization that has explicitly opted in to that use, or for a design partner. No other organization's retained payloads are used this way, and this use is never on by default. The export is not a scheduled job and nothing in the gateway calls it automatically; an engineer runs it by hand, against the opted-in or design-partner orgs only. Turning payload logging on is a separate setting from that opt-in and is not a self-serve dashboard toggle today: contact us to enable it for your organization. For an org that has opted in, the export reads only rows that org already retained under payload logging, including any payload you asked us to keep for debugging: rows written while ZDR is on, or while payload logging is off, carry no prompt or completion text, so the export has nothing to read for your org in either state. There is currently no per-request control that lets an opted-in org exclude a specific retained row from the export; not opting in is the org-level control. Staying on the metadata-only default (or asking us to turn payload logging back off) and Zero Data Retention above stop content from being written in the first place, independent of the opt-in setting. If keeping your prompts out of any internal use is a requirement, do not opt in, and stay on the metadata-only default or run with Zero Data Retention on.

PII redactor inline path

Status: Pilot, by request. The gateway PII redactor runs at the request boundary in production today. It is off by default for every org and protects nothing until it is turned on. Turning it on is not a self-serve dashboard toggle today: contact us to enable it for your organization. Once enabled, the two settings below control how it behaves.

The PII redactor runs inline on the request path so an org with the inline flag enabled redacts the prompt before it leaves the gateway, then re-injects the original spans on the response stream so the model's reply still references the user's actual data.

The two flags

The redactor reads two per-org settings:

  • A redaction mode: one of off, redact_in_transit, or redact_and_log. off skips the redactor entirely.
  • An inline-enabled gate. The inline path described here runs when the mode is not off AND inline is enabled. With the mode set but inline still off, the older write-through path runs instead: it redacts with [REDACTED:TYPE] markers and does not re-inject.

The two-flag design lets an org stage the rollout: set the mode to redact_and_log first to confirm the redactor finds the right spans without changing the prompt that goes upstream, then enable inline once the audit trail looks correct.

The two modes

redact_in_transit rewrites the prompt to replace each PII span with a placeholder token (<PII:EMAIL_0>, <PII:PHONE_1>, <PII:SSN_2>), stashes the original spans server-side with a 60-second TTL, and re-injects the original spans into the streaming response so the user sees their actual data in the model's reply. The model never sees the PII; the user never sees the placeholders.

redact_and_log rewrites the prompt the same way but skips the re-injection. The user sees the placeholders in the response. This mode is for high-sensitivity workflows where the PII must not round-trip even into the response stream.

The 60-second TTL

The span map lives for exactly 60 seconds. A redactor run that takes longer than 60 seconds loses the spans and the response surfaces the placeholder tokens instead of the original data. The 60s budget is generous for the typical model latency (single-digit seconds) but tight enough that a stalled response cannot leak spans into a future request's window.

The request id that keys the span map is server-generated and is never read from a client-controlled header, so an attacker cannot forge one to read another user's span map.

The detection pass

The redactor's detection today is four regex patterns: email addresses, phone numbers, Social Security numbers, and credit card numbers. There is no external model behind it: the redactor does not call another service, and nothing beyond those four pattern types is recognised, DPI setting included (see below).

The latency-budget gate

A per-request wall-clock budget (default 800ms) caps the redactor's time regardless of prompt length. On timeout, the redactor fails closed for every org by default: the request is rejected (HTTP 422) rather than forwarded unredacted, and the gateway writes an audit row carrying the entity count and types only (never the unredacted content).

Fail-closed is the safer of the two failure modes for a feature whose job is keeping PII out of what leaves the gateway: it never lets unredacted content through, at the cost of an occasional rejected request under sustained latency pressure.

On the roadmap

  • A per-team override on the redaction mode so a team can be more aggressive than the org default without forcing every team to match.
  • A live latency p95 plus p99 readout on the dashboard tile so an admin can spot a slow cold start before it triggers the budget gate.

Anomaly rate-limit

Status: Live today. The per-org anomaly rate-limit and its anomaly_rate_threshold alert-rule are in force in production.

The anomaly detector emits one event for every detection. A misconfigured detector or a real spike can flood the store with thousands of events per minute, which both costs storage and drowns the alert-rule evaluator. The per-org rate-limit caps the emission rate.

The per-org cap

A per-org rate limit (default 10 per minute) caps how many anomaly events the detector emits per minute, counted separately for each anomaly type so one noisy heuristic cannot starve the others. The value cannot be set below 1, so a misconfigured 0 cannot silence the detector entirely.

When the cap is hit, the detector still writes the event but flips a rate_limited flag to true so the dashboard can render a "Rate limited" badge on the row. The row write is the audit evidence that the detector fired; what the flag suppresses is the webhook notification, so the audit trail stays complete without fanning every downstream subscriber out on a runaway detector.

The alert-rule type

An anomaly_rate_threshold alert-rule fires when the count of anomaly events in a rolling window exceeds a configured threshold. The rule config carries two keys:

{
  "window_hours": 1,
  "threshold":   50
}

window_hours defaults to 1 and threshold defaults to 50. The dashboard form constrains the config shape to exactly those two keys; an attempt to inject any other key is rejected at the submit handler.

The cooldown

Every alert-rule carries a cooldown that prevents the same rule from firing twice in quick succession. The dashboard prefills a 15-minute cooldown for the anomaly_rate_threshold rule type; a create that omits the field entirely lands the 60-minute default the other rule types use. The floor is enforced server-side, so even if the form submits a zero cooldown the validator rejects it and a misconfigured rule cannot ladder.

On the roadmap

  • A per-rule plus per-recipient cooldown so the same rule can fire immediately to the on-call SIEM but every 60 minutes to the Slack channel.
  • An "anomaly digest" alert-rule type that batches the rate-limited rows into a single notification at the cooldown boundary so the alert recipient sees one rolled-up message instead of N rate-limited badges in the dashboard.

Access scoping

Status: Live today. Org-scoped enumeration, operator gating, and an owner/admin role gate on sensitive admin actions are all in force in production. The granular multi-permission catalogue described under "On the roadmap" below is not wired to any route yet.

Every admin surface enforces org-scoped enumeration. A small set of platform-operator surfaces (the cross-org and fleet-wide views) gate on an operator flag. Separately, an owner/admin role check already backs specific sensitive actions across most of the admin surface, including key reveal, billing, and member management. The granular multi-permission catalogue, which would name permissions individually per role instead of today's owner-or-not check, is on the roadmap.

Org-scoped enumeration

Every admin route resolves the caller's active org and masks the result set to that org. An admin with credentials for org A cannot enumerate, read, or mutate org B's keys, audit rows, billing, or settings; the query is scoped to the caller's org before it runs, so a cross-org probe returns the caller's own data or a 404, never another org's. This is the load-bearing tenancy boundary and it is enforced on every admin handler.

Operator gating

A small set of platform-operator surfaces (the cross-org and fleet-wide views) gate on an operator flag carried on the session. A non-operator session that scripts the endpoint directly is masked with a 404 before the handler runs, so the gate is enforced server-side rather than just hidden in the UI, and the route stays indistinguishable from one that does not exist.

Role assignments

Every member carries an org-scoped role of owner, admin, or member, set on invite and changeable from the Settings → Members pane by an owner or admin (a plain member sees the role read-only). Promoting someone to owner is restricted to existing owners, and the last remaining owner cannot be demoted or removed. Every invite, role change, and removal writes an audit row. This role is a real enforcement input today: it is what gates key reveal, billing, and member management above, not a display-only label.

The Members pane also surfaces a separate, read-only "Roles" panel backed by the granular multi-permission catalogue on the roadmap below. That catalogue is not yet wired to any route guard, so the panel does not drive access; the role in the Members list above is the boundary the gateway actually enforces.

On the roadmap

  • The granular multi-permission catalogue (named permissions such as KEY_REVEAL, AUDIT_VERIFY, and BILLING_READ, mapped per role) wired to a uniform per-route guard, replacing the hand-rolled owner/admin check most admin routes carry today.
  • The read-only Roles panel wired to that catalogue's live enforcement state, once the guard above lands, instead of the separate table it reads from today.

Multi-region scaffolding

Status: Roadmap. The data-residency control surface records an org's intended region today, but provider-pinned data-residency (routing an eu org only to eu-region providers) is on the roadmap and not in force yet.

The current build lands the data-residency control surface plus the read-side resolution so a later release can wire region-pinned providers without re-architecting. The region currently records the org's intended residency; provider pinning is on the roadmap.

The region values

An org's data-residency region is one of three values: us, eu, or apac. The default is us, which keeps every existing org on the US region until an admin sets the value via the Settings → Security pane. A fourth value is rejected, so a typo cannot land a rogue region.

Defaulting

The region resolution defaults to us whenever the stored value is missing or empty, so the read path is safe to call across a deploy window before every org has a region set.

On the roadmap

  • Region-pinned routing so an eu org only routes to providers whose region matches.
  • A per-team override on the region so a team in a US-default org can opt into EU residency without forcing every team to match.
  • A region-aware webhook deliverer so a US-default org with a Frankfurt receiver routes the delivery through the EU egress pool rather than the US pool.
  • A region-pinned audit-log composer so an evidence pack only carries rows from the region the auditor is scoped to.

PII customer-scoped enablement

Status: Pilot, by request, except DPI (see below). These controls are part of the gateway PII redactor, which runs at the request boundary today and is off until IQ Routing enables it for your organization on request.

The redactor's latency budget, the deep-PII inspection (DPI) toggle, and the re-injection policy are per-org controls so an org can opt into a wider budget plus a tighter re-injection policy without a process restart. The inline path reads the current values on every request, so a settings change lands on the very next request. The DPI toggle is stored the same way but does not change redactor behaviour today; see DPI status, below.

The three controls

  • A per-org latency budget (default 800ms): the wall-clock bound the inline path hands the redactor.
  • A per-org DPI opt-in (default off): reserved for a deeper machine-learned recognizer. See DPI status, below, before enabling it.
  • A per-org re-injection policy (default always): controls how the response wrapper re-injects the original spans.

The latency budget override

An org that pegs the budget at 100ms gets its request rejected (HTTP 422) on every prompt that exceeds 100ms; an org that widens to 3000ms gives the redactor more headroom before it times out. The dashboard clamps the input between 50ms and 3000ms to protect the operator UI from typos and runaway values. A submission outside the accepted range is rejected at the validator boundary.

DPI status

The DPI toggle exists in the org settings and in this API, but the transformer-backed recognizer it is meant to load has not shipped: the current deployment does not have the model available to load. Turning DPI on does not change detection today -- the redactor still runs the same four regex patterns described above, with no additional coverage and no cold-start latency. Leave DPI off until this page says otherwise; it is not a live wider-coverage control yet.

The re-injection policy

The re-injection policy takes one of three values: always, trusted_only, or never. Any other value is rejected so a misconfigured admin endpoint cannot land a rogue policy.

always is the default and matches the inline contract: the response wrapper re-injects every per-request PII span so the user sees their original data in the model's reply. trusted_only re-injects only when the response originated from a trusted endpoint; otherwise the placeholders stay in the response. never skips the re-injection entirely so the response stream carries the placeholders.

The Settings → Security pane exposes the three values and carries an inline caption naming the default-on contract.

Anomaly detector extended heuristics

Status: Live today. The base heuristic, the five extended heuristics, and the per-type rate-limit windows are in force in production.

The anomaly detector fires one base heuristic (thinking_budget_excess) plus emits one event per detection. The detector is extended with five new heuristics covering request-rate spikes, cost-rate spikes, failed-auth bursts, key-vault-reveal bursts, plus webhook failure streaks. Each event carries a discrete anomaly type so the dashboard can index-seek per type instead of scanning a free-form reason field.

The five new types

The five new heuristics each emit an event tagged with a discrete anomaly type, and that type is the label you see on the dashboard row and in the anomaly-event telemetry:

  • request_rate_spike: fires on a sudden jump in request volume against your org's own recent baseline, so a new or low-traffic team does not false-flag on its first normal burst.
  • cost_rate_spike: fires on a sudden jump in spend rate against your org's own recent baseline, the same shape as request_rate_spike but measured on cost instead of request count.
  • failed_auth_burst: fires when a single actor's failed-auth attempts spike in a short window, independent of your org's traffic baseline, since a burst of auth failures from one actor is anomalous regardless of how much traffic your org normally sees.
  • key_vault_reveal_burst: fires when a single actor's key-reveal actions spike in a short window. The bar for this type is tighter than failed_auth_burst, since legitimate key reveals are rare.
  • webhook_failure_streak: fires when a single webhook destination accumulates consecutive delivery failures with no successes in between. See webhook capacity planning for how to keep a receiver healthy enough to avoid tripping this.

The exact thresholds behind each heuristic are tuned internally and are not published, since they double as anti-abuse controls; treat the anomaly type as the stable, documented surface, not the number that triggers it.

Per-type rate-limit windows

The per-org rate-limit is scoped to the firing anomaly type rather than shared across every heuristic. The detector counts the events already written for that org, that team, and that anomaly type inside the last rolling minute, and compares the count against the org's cap; a detection that would exhaust the type's own bucket flips the rate_limited flag on the emitted event but writes the event anyway, so the audit trail stays complete either way.

Per-type scoping keeps the budgets independent: each anomaly type gets the full per-minute cap to itself, so a noisy heuristic on one type cannot suppress visibility into another, and tuning the cap up for a high-traffic org does not trade one type's visibility against another's.

Alert-rule integration

The anomaly_rate_threshold alert-rule counts every anomaly event the org wrote inside the rolling window, without filtering by type, and fires when the count exceeds the configured threshold. Its config is a closed schema of exactly window_hours plus threshold; a third key, including an anomaly_type filter, is rejected at the gateway boundary rather than silently stored.

One heuristic does have its own rule type. auth_failure_burst_5m reads failed_auth_burst events over a fixed five-minute window, which is the shape a credential-stuffing spike actually has, so a login-loop or a stuffing run can page separately from the org-wide anomaly count.

The 15-minute cooldown prefill carries forward on the anomaly_rate_threshold rule type; auth_failure_burst_5m takes the 60-minute default unless the operator dials it down.

On the roadmap

  • A per-detector backoff so a heuristic that keeps firing and hitting its rate limit tightens its own budget for the next window, closing the gap a patient, low-and-slow actor could otherwise pace around.
  • A per-heuristic alert digest that batches the rate-limited events into a single notification at the cooldown boundary.
  • Full audit-chain coverage on the anomaly event records so the verifier no longer skips any rows.