Skip to content

OpenAI-Compatible Model Gateway

Editions: OSS. Contributor documentation for this repository.

Managed runtimes send model traffic through Preloop's OpenAI- and Anthropic-compatible gateway. This chapter covers gateway ingress, accounting, budgets, and the runtime session identity layer used for browsing and audit.

Dedicated gateway processes install plugin request dependencies and run on_gateway_startup / on_gateway_shutdown hooks. A required gateway startup hook failure prevents serving traffic. API routes and API-only background workers remain in the api / all process. Plugins without gateway-specific resources inherit no-op lifecycle hooks.

OpenAI-Compatible Model Gateway

  • Purpose: Centralize model traffic from managed runtimes behind Preloop control.
  • Ingress: GET /openai/v1/models, POST /openai/v1/chat/completions, POST /openai/v1/responses, POST /openai/v1/embeddings
  • Additional Client Compatibility: POST /anthropic/v1/messages for Anthropic-format clients such as Claude Code
  • Subscription-OAuth Passthrough: Anthropic-protocol requests backed by a Claude Code subscription-OAuth credential bypass the LiteLLM transcode and are forwarded to Anthropic verbatim (system block array, cache_control markers, and client anthropic-beta flags preserved; merged with the OAuth beta flag). The upstream validates the structural shape of subscription-OAuth requests, and the transcode would destroy it. Budget preflight, governance tool-stripping, attribution, and usage accounting still run on this branch; API-key and OpenAI-protocol traffic keep the LiteLLM path.
  • Native Responses Passthrough: POST /openai/v1/responses forwards to the upstream's own /responses endpoint whenever the upstream is OpenAI-shaped (LiteLLM openai/<id>: BYOK OpenAI, openai-compatible, custom endpoints) and authenticates with an API key, instead of transcoding the request into a chat-completions call. The transcode dropped instructions, reasoning, include, store and prompt_cache_key, forced Codex tool shapes through codex_tool_compat flattening, and could not work at all against upstreams that implement Responses but not chat completions (measured: OpenCode Zen). Non-streaming responses and SSE frames are relayed verbatim, with a side parser collecting response id, usage and assistant text for accounting. Upstreams with no Responses endpoint are detected from a 404/405/501 and remembered per base URL (15 min TTL), so chat-completions-only deployments fall back to the transcode with no configuration; meta_data.gateway.responses_api (auto default, native, transcode) pins the choice. OpenRouter stays on its LiteLLM adapter because that adapter carries usage accounting and app attribution. Budget preflight, governance tool-stripping, attribution, response policy and usage accounting all run on this branch; message-level context optimization does not, for the same payload-fidelity reason as the Anthropic passthrough.
  • Embeddings: POST /openai/v1/embeddings takes the standard OpenAI embeddings body (model, input, optional dimensions / encoding_format / user), resolves and authorizes the model through the same account-scoped alias resolution as the completions routes, and returns the upstream vectors unchanged. Kill switch, budget preflight, retry ownership and usage accounting all run on this branch, so one embeddings call lands one ApiUsage row with its account, key, model, prompt tokens and catalog-priced cost; an embedding model missing from the price catalog records cost_source='unpriced' with a NULL cost, never a silent $0. There is no streaming variant and no tool handling. API-key and ambient (Bedrock) credentials are supported; subscription OAuth is not, because no subscription upstream exposes an embeddings endpoint. The stored ledger/event copy of the response keeps the model, usage and the count and width of the returned vectors, not the float arrays themselves.
  • OpenCode Client Identity: Native Responses forwarding to the verified OpenCode Zen endpoint preserves a bounded allowlist of the caller's actual User-Agent and x-opencode-* identity headers. It never invents missing OpenCode identity or forwards the caller's credentials, cookies or hop-by-hop headers. Upstream authentication remains model-specific. Other destinations retain Preloop attribution. This preserves the original client identity for provider eligibility checks; it does not override the provider's free-tier policy.
  • Hosted OpenCode Protocol Selection: Gateway models retain their API protocol in OpenCode's per-model provider registry, including primary work and title generation. meta_data.gateway.responses_api=native selects the Responses adapter for upstreams eligible for native gateway forwarding; transcode retains chat completions. Auto mode includes an endpoint-scoped snapshot of OpenCode Zen Responses models from models.dev. The canonical endpoint, source, snapshot date and 90-day review interval are named constants in model_api_protocol.py. A process logs one maintenance warning when auto-selection consults an overdue snapshot; age never changes routing, and CI uses deterministic dates without fetching the live catalog. Refresh by reviewing the opencode provider, selecting models whose effective SDK is @ai-sdk/openai, and updating the model set and snapshot date together. Unknown models keep the chat adapter, so one authorized inventory can contain both protocols. Explicit protocol overrides bypass the snapshot and support new models or separately verified endpoints; endpoint mirrors are not assumed to share capabilities. CLI onboarding and direct-provider OpenCode configuration retain their existing behavior.
  • Retry Ownership and Error Reporting: The gateway owns a bounded retry budget for transient provider failures; nested LiteLLM and SDK retries are disabled per gateway call. Connection failures, timeouts, overload and opaque upstream 5xx remain eligible for recovery. Explicit protocol incompatibility is a terminal upstream_protocol HTTP 400, and native Responses retains raw failure status and Retry-After until final normalization. A hosted model the deployment never gave an operator tariff is a terminal hosted_tariff_unconfigured HTTP 503: it is 5xx-shaped but nothing upstream was contacted, so the gateway attempts it once, the usage row records the configuration class rather than upstream_error, and agent harnesses read it as a configuration failure instead of a provider outage worth resuming. Once streaming output has been emitted, the gateway never replays the request; existing flow reconnect/resume behavior is unchanged. Automatic Sentry OpenAI SDK captures are filtered only for recognized upstream status/connection exceptions inside a gateway-owned attempt. Final classified failures remain in gateway usage/events and admin alerts; unowned SDK calls and local application errors remain reportable.
  • DeepSeek Responses Reasoning Continuity: For resolved deepseek provider rows using chat transcode, provider reasoning_content is carried as an opaque Responses reasoning item with an empty summary and Fernet-encrypted encrypted_content. Both complete responses and SSE output-item events preserve it. Replayed items are bound to the authenticated account, resolved model/configuration, and stable tool-call IDs, so Codex namespace/freeform translations do not break continuity. The gateway reconstructs the assistant tool turn and restores the original provider field. Encryption uses the existing shared deployment key and works across gateway replicas without an in-process history cache. Encrypted items have a 4 MiB ASCII envelope limit, checked before decryption, and no short expiry that would break delayed continuation. Original assistant text is checked when reconstructing a tool turn. Malformed or cross-scope items fail with invalid_reasoning_content; they never downgrade to a fallback. Only legacy DeepSeek assistant turns without preserved reasoning receive the provider-compatible empty field, which cannot recover reasoning already lost by an older gateway. Existing supplied content is retained. Native Responses passthrough and other providers keep their existing behavior. Clients must replay the opaque output items alongside tool results, following the Responses reasoning contract.
  • Claude Model-Family Fidelity: Claude Code onboarding imports one gateway model per selectable family (opus/fable/sonnet/haiku) sharing a single credential secret. By default it does not write the stock ANTHROPIC_DEFAULT_OPUS_MODEL / _SONNET_MODEL / _HAIKU_MODEL pins, so /model selectors follow Claude Code's own built-in defaults and, with subscription OAuth and family autoregistration enabled, a new Anthropic release arrives with the next Claude Code update; only the Fable pair stays pinned because stock Claude Code has no built-in fable default. /model switching, background fast-path requests, and subagents keep native UX while routing through Preloop. preloop agents onboard --pin-model-families (and refresh --pin-model-families) restores explicit per-family pins for API-key accounts or gateways where autoregistration is disabled; API-key accounts otherwise need preloop models sync to populate new identifiers before selecting them. The choice is persisted in the local enrollment state so a later flag-less refresh honours it. Context-window variant markers (claude-fable-5[1m]) resolve/price against the base model but are forwarded verbatim upstream on the OAuth passthrough. Unknown claude-* identifiers requested over a subscription-OAuth credential (e.g. new dated snapshots after a Claude Code update) are lazily auto-registered against the same credential and bound to the requesting managed agent (model_gateway_claude_family_autoregister_enabled, default on); Anthropic remains the authorization boundary for what the subscription may use, and subject-scoped allowed_models checks still apply. With explicit pins enabled, refresh verifies candidate family models against the live Anthropic list before upgrading; an authorized current pin is retained when verification is unavailable or the candidate is absent. Fable uses the same verified resolution. Before a row is minted, the identifier is optionally verified against GET /v1/models/{identifier} with the template's OAuth token (model_gateway_claude_family_autoregister_verify_upstream, default on): a 404 blocks registration so a typo or a guessed snapshot date cannot become a permanent catalog row, a 200 registers and marks the row upstream_verification=verified, and any other status or transport error registers as before, marked unverified, with one warning. Positive results are cached for 24h and 404s for 10min; set the flag false to skip the extra call entirely.
  • Codex ChatGPT-OAuth Model Fidelity: Codex CLI onboarding imports the models present at onboard time onto openai-codex rows sharing a single ChatGPT OAuth secret. Codex's /model picker is built into the CLI, so a newly shipped identifier (e.g. gpt-6-astra) can be selected before the Preloop catalog knows it. preloop models sync cannot list against those credentials. Unknown gpt-* / o-series / chatgpt-* identifiers requested over a Codex subscription-OAuth credential are lazily auto-registered against the same secret and bound to the requesting managed agent (model_gateway_codex_family_autoregister_enabled, default on); OpenAI remains the authorization boundary, and subject-scoped allowed_models checks still apply.
  • Streaming: Supports SSE streaming for chat completions and responses.
  • Authentication: Reuses short-lived runtime bearer tokens while preserving ApiKey context and runtime-principal metadata.
  • Database connection lifetime: HTTP bearer authentication runs in one database worker and releases its owned session after preserving scalar auth state. Credentials and revocation are still checked on every request. Request and response policy reads also release their lookup transactions before provider or approval waits. The streaming buffering gate reloads current rules at the first stream pull, including rules added during request approval or provider setup. Required final checks on buffered output reload current policy again. An unavailable initial stream check closes its upstream resource, sends a protocol-specific error without model output and records a failed request. Internal callers retain ownership of their sessions.
  • Accounting: Persists token usage and estimated cost in ApiUsage, including first-class cache-read/cache-creation/reasoning token columns, currency, and provenance markers (cost_source: override | model_config | provider | catalog | subscription | reconciled | unpriced; usage_source: provider | estimated | partial). When the upstream reports the request's actual cost in its usage payload (OpenRouter usage accounting: usage.cost / usage.cost_details.upstream_inference_cost; the gateway requests it via usage: {"include": true} on OpenRouter-bound requests), that figure is authoritative over catalog estimates and the row is tagged cost_source='provider'. Historical rows that predate usage accounting can be repriced from the provider's daily activity ledger via scripts/backfill_openrouter_ledger.py (OPENROUTER_ACTIVITY_KEY env var, dry-run by default): each day's ledger total is allocated across that day's unpriced rows proportionally by tokens and tagged cost_source='reconciled' with an audit marker in meta_data: labeled approximations, deliberately distinct from per-request provider figures. Streaming requests always request the provider's final usage chunk (stream_options.include_usage is injected upstream; the synthetic chunk is stripped from clients that did not opt in); usage payloads split across chunks are merged, client disconnects record partial usage at status 499, and a local tokenizer fallback (litellm.token_counter) estimates tokens when a provider reports none.
  • Budget rollup transaction: Priced usage and every account, subject, model, and period spend bucket commit together. The CRUD layer performs one ordered PostgreSQL batch upsert and leaves the commit to usage recording, so a failure cannot publish a partial set of counters. This reduces transaction overhead without changing preflight budget checks: concurrent requests can still pass a check before their eventual spend is recorded. Atomic reservation and retry deduplication are separate guarantees.
  • Budget Enforcement: Applies account-level, flow-level, and subject-scoped allowed-model checks before upstream dispatch. Preflight input estimation uses a real tokenizer with a chars/4 fallback; OAuth-subscription-credentialed models preflight at $0.
  • Pricing: Default prices come from a vendored, provenance-stamped snapshot of litellm's price map (services/data/model_prices.json, registered via litellm.register_model at startup) so estimates are deterministic per release. The snapshot is kept small on purpose: Preloop-routed providers only, modes the gateway serves only (chat, responses and embeddings; no image/audio/video models, and embedding rows without a per-token input price are left out so a per-query multimodal model records as unpriced rather than $0), past-deprecation entries dropped, entries priced only through tiered_pricing vendored at their lowest published tier so a tier-only row is not stored priceless, fields stripped to pricing essentials; scripts/update_model_prices.py is the standardized refresh path (fetch → review diff → commit, with a --check CI staleness gate). Alibaba Cloud Model Studio USD sites use a Singapore International seed plus a live native GET /api/v1/models overlay (alibaba_pricing / alibaba_price_catalog) instead of the LiteLLM dashscope/ map. The snapshot only changes on deploy, so every process also refreshes the upstream map at runtime (model_price_catalog.PriceMapRefresher, started with the app and the NATS worker): on startup and every MODEL_PRICE_MAP_TTL_SECONDS (~6h) it downloads the map and registers new and changed prices over the snapshot, skipping operator-reviewed prices (preloop_price_provenance from the reviewed feed) and the snapshot's first-party Moonshot/z.ai overlay rows. Models still missing self-heal on a miss: an unpriced usage row, or a preflight denial for a model without a price, triggers a one-shot background lookup against the same cached map (model_price_catalog.schedule_price_lookup) that registers the price and re-prices the triggering row; failed downloads back off (MODEL_PRICE_MAP_FAILURE_BACKOFF_SECONDS, ~15m), and unknown model names are negative-cached (~24h) so the same model never triggers repeated lookups. model_price_live_lookup_enabled gates both paths. Each fetch logs one line (Model price map fetch: source=... outcome=..., with entry count and sha256 on success) and each negative-cache insertion logs one WARNING naming the alias by namespace and hashed token (raw model names are never logged; compare against _model_log_token(<alias>)); /health reports the last fetch time, outcome and entry count under model_price_map. Model-name resolution normalizes Bedrock region prefixes (us./eu./…) and trailing date stamps. Account-scoped price overrides win over the catalog (resolved once, via services/pricing_overrides.py, for the record path, budget preflight, execution metrics, and tool stats) and support per-token-type prices, fixed request fees, discounts, prepaid balances, effective-date ranges, and non-USD currencies via an explicit fx_rate_to_usd (all stored costs remain USD). Requests routed over subscription OAuth credentials (Claude Code Max, ChatGPT/Codex) record $0 spend with the API-equivalent value kept in metadata. Unknown models record NULL cost tagged unpriced and are surfaced in the Cost overview; the Enterprise reprice endpoint (POST /api/v1/billing/cost/reprice) re-derives historical costs from stored tokens against current prices (analytics-only; budget spend is never rewritten).
  • Rate-Limit Telemetry: Provider rate-limit headers observed on gateway upstream responses (Retry-After, anthropic-ratelimit-*, x-ratelimit-*) are parsed by services/rate_limit_telemetry.py and persisted on ApiUsage (rate_limit_retry_after_ms column plus a verbatim header snapshot in meta_data["rate_limit"]). 429s are subtyped transient vs quota-exhausted (delegating to the shared upstream-error taxonomy when available, with a labeled heuristic fallback). GET /account/gateway-usage/rate-limits aggregates 429 counts, provider-advised blocked time, and the latest per-model headroom snapshots for the Console's Rate Limits & Headroom panel; every reported number echoes an observed provider response, nothing is estimated.
  • Admin Alert Throttle: Individual upstream disconnects are logged and returned as gateway/SSE errors without emailing admins. Other final 5xx failures reserve one notification per quiet window in NATS JetStream KV. General availability failures from the same public provider and canonical endpoint share a budget across accounts, models, failure classes and upstream 5xx statuses. The alert describes a representative failure; detailed usage/error attribution remains unchanged. Credential, quota and rate-limit failures retain granular incident keys; private/local endpoints and endpoint-less generic adapters keep account/model isolation. Keys are hashed and omit exception messages and credential-bearing URL parts. GATEWAY_ERROR_ALERT_INTERVAL_SECONDS sets the window (default 300; finite positive seconds, otherwise the default). Each process applies the same budget locally before queueing, bounding broker work during broad outages. Suppression counts cover only that process's previous local window, not fleet totals. The bounded notification worker (32 queued/running notifications) performs reservation and delivery, so requests never wait on NATS or notifications. A dedicated running event loop owns a reusable NATS connection and current KV handle, keeps heartbeats active while idle, and closes stale clients before reconnecting on a later eligible alert. Reservations have a one-second deadline and separately bounded connection cleanup; shutdown rejects late work and closes the client before stopping the loop. NATS atomically creates the key; broker TTL expires it without relying on gateway clocks. Buckets are separated by interval configuration, retain one revision, and cap storage at 4 MiB per bucket. A broker outage or full bucket falls back to the reserved local budget, temporarily permitting one alert per process per notification budget/window. Local state caps active budgets at 4096; saturation drops new ones until slots expire. Failed delivery or a full queue consumes its reservation to avoid notification retry storms.
  • Context Optimization: Subject-scoped request-context optimization on the hot path (services/context_optimization.py): repeated-prefix dedupe, noise stripping, and tool-result caps applied before upstream dispatch, with evidence-grounded savings attribution.
  • Observability: Emits normalized model-call events with redaction-aware request/response payload capture, provider-neutral conversation previews (up to MODEL_GATEWAY_MAX_PREVIEW_CHARS, default 32768), prompt-cache token breakdowns, and optional indexing into a gateway search corpus (MODEL_GATEWAY_AUTO_INDEX_INTERACTIONS). Session content is also chunked into a session-scoped search corpus as each source row is written (gateway interaction, transcript message, tool call, operator note, session summary; flow logs have a writer but no call site yet). Stored chunk text is gated by MODEL_GATEWAY_CAPTURE_CONTENT; SESSION_SEARCH_INDEX_ENABLED disables the session corpus writes. Optional OTLP export (disabled by default) emits OpenTelemetry GenAI spans whose gen_ai.conversation.id matches runtime session identity; see docs/guide/observability-otlp.md. Export errors are logged and never fail the user-facing request.
  • Debug Surface: Flow execution-scoped gateway events can already be queried via the flows API, runtime-session explorers can query recent session activity directly, and the operator dashboard can aggregate active sessions, recent tool calls, and daily model spend.
  • Managed Agent Onboarding: External agents such as OpenClaw can be enrolled so local model traffic is rewritten onto this gateway while local MCP configuration is narrowed to Preloop-managed proxy access.

Runtime Session Identity

  • Purpose: Provide a shared identity layer for browsing, auditing, and searching managed runtime sessions across both flows and onboarded external agents.
  • Current Implementation: A new additive RuntimeSession layer now uses flows as the first session source, bridged through runtime_session_id, flow_execution_id, runtime-principal metadata, and optional agent_session_reference.
  • Session search corpus: Writes land next to each source row. A session title or summary persisted by the plugin title hook, usage-import metadata, or a gateway auto-summary writes one session_summary chunk. Content capture gates stored text; SESSION_SEARCH_INDEX_ENABLED stops new writes without a restart. Clearing a title still drops its stale chunk so the corpus cannot keep answering with text the session no longer carries.
  • Current Explorer Surface: Account-scoped runtime session list/detail endpoints now expose recent managed sessions plus their captured gateway interactions so the console can drill from aggregate usage into one session timeline. GET /account/runtime-sessions/{id}/requests reads per-request ApiUsage rows (tokens, cost, status, tool-schema attribution) to power a unified session replay with turn/delta deduplication, sortable chat view (newest-first by default, message order following the turn sort), cache-token visibility with the re-sent prompt-cached prefix collapsed inside full request context, and inline operator activity turns (session-replay-panel, preloop-session-observer in the Console). Operator ergonomics on top: keyboard navigation over turns (j/k/arrows, Home/End, Enter/o to expand), clickable summary-bar stats that jump to the most-expensive or first-failed turn, relative turn timestamps with absolute time on hover, a session list that collapses into a compact picker bar (animated, reduced-motion aware) once a session is chosen, and a deep-linkable replay mode via ?replay= where the host view opts in (syncModeToUrl).
  • Reporting Query Bounds: Session lists combine account-scoped direct and legacy flow attribution with disjoint UNION ALL branches and equality joins; a request matching both identities counts once per session. Aggregation still precedes pagination to preserve usage-based ordering. Latest model/provider lookup returns one database row per displayed session. Tool-cost breakdowns project only pricing/principal fields and the tools_meta JSON subtree, retaining the existing newest-5,000-row reporting limit without loading unrelated request or response metadata. These reporting reads stay in the CRUD layer.
  • Automatic Gateway Session Summaries: The gateway optionally summarizes successful model activity using the account's default model. Failed primary requests skip this work. Summary attempts occur after the first successful call and each tenth successful call, even when an earlier summary attempt failed, so a missing summary does not cause generation on every request. The count comes from the shared CRUD layer; this cadence is not an atomic lock across concurrent requests. Optional generation uses a separate service instance to isolate retry, rate-limit and caller-identity state. Its failures remain diagnostic summary logs without main-gateway admin alerts or changes to the primary result. A successful persist also reconciles that session's session_summary search chunk, and a failure there never rolls back the summary row. An identical regeneration leaves summary_updated_at and the search chunk unchanged, matching the plugin title hook and usage-import metadata writers.
  • Summaries & Titles: Opt-in LLM-generated session summaries (POST /account/runtime-sessions/{id}/summaries) and background session titles use the plugin service registry. Session-list response callbacks only submit scalar IDs to session_title_scheduler; they do not open a DB session or wait for a model call. The plugin owns bounded admission, deduplication, worker DB lifetimes, and shutdown. Missing or failing optional schedulers leave titles for a later list load. Enterprise billing retains per-account daily spend caps (billing_session_optimization_daily_cap_usd, billing_session_title_daily_cap_usd).
  • Auxiliary Model Credentials & Fallback: Non-interactive auxiliary generations (approval summaries, session titles, policy generation, issue compliance/duplicates/dependencies) resolve model credentials through the secret service via services/model_credentials.resolve_model_call_credentials, so vault-backed (credentials_secret_id), legacy plaintext, OAuth, and ambient credentials all work; routing (api_base) is preserved even when credential resolution fails. Approval summaries and session titles additionally retry once against the system-wide default model when the account's model fails, bounded by a per-account, per-UTC-day cap (PRELOOP_AUX_FALLBACK_DAILY_CAP, default 50, enforced per process). The main gateway/completion path never falls back.
  • Session Identity On The Wire (public semantics): A gateway request is bound to a conversation by the FIRST of these that is present, highest precedence first:
    1. X-Preloop-Session-Id: Preloop's own explicit override. Always wins, works on every gateway ingress (OpenAI, Anthropic, Gemini), and is the documented answer for any client that wants deterministic control. On a plain API key (one whose context_data carries no runtime_principal and no pinned runtime_session_id) it is also the per-request opt-in that binds otherwise sessionless traffic to a runtime session keyed session_source_type="api_key" / session_source_id="{api_key_id}:{client_session_id}". The account and key id are part of that key, so a caller-supplied id can never adopt another key's or account's session. No header, or an invalid/oversized value, records usage with runtime_session_id NULL exactly as before; at most one session is created per (account, key, id).
    2. A vendor-native session signal the agent already sends, read only when the credential's runtime_principal.type identifies that agent: Claude Code's X-Claude-Code-Session-Id / metadata.user_id, Codex's Session-Id / Thread-Id, OpenCode's X-Session-Id. These are deliberately gated on the principal type because Session-Id and X-Session-Id are generic names that any intermediate proxy, load balancer, or CDN may stamp; a wrong session boundary is unrecoverable after the fact, since boundaries can never be re-derived from stored rows. A plain API key has no runtime_principal.type, so none of these are trusted for it and a vendor header never opts a plain key in.
    3. prompt_cache_key on the OpenAI-shaped ingress. OpenAI splits the old user field into safety_identifier (stable principal) and prompt_cache_key (per-conversation, and explicitly "replaces the user field"). Agents populate it for their own cache hit rate, which makes it a de-facto convergence point. It ranks below the two above because it is a cache key, not an identity key: a client may legitimately share it across conversations with identical prefixes or rotate it on compaction. On a plain key it is likewise ignored for session identity, because the plain-key opt-in is explicit and header-only.
    4. Nothing. Signal-less sources (Gemini CLI, Hermes, OpenClaw's Anthropic transport) fall through to the inactivity closer below. Anything read here is normalized, and a hostile or unusable value degrades to source-keying rather than being trusted. Session drill-down, Optimize and replay read the per-request ApiUsage rows attached to the session, so a plain-key session becomes usable for them only once it has attributed requests; a session with no captured requests stays normally replay-ineligible rather than reporting fabricated savings.
  • Parent session lineage: Runtime sessions carry a nullable parent_session_id when a harness states who spawned the turn on the wire. Two harnesses do that today: OpenCode sends X-Parent-Session-Id next to X-Session-Id; Claude Code keeps sending the parent's X-Claude-Code-Session-Id and adds X-Claude-Code-Agent-Id, so the child is keyed <session>:<agent> and the parent is the session the header already names. Hook ingest maps parent_conversation_id the same way. Chat, responses, and embeddings all pass the derived parent; an operator X-Preloop-Session-Id wins the session key and records no parent. Lineage is write-once: stored on create, a NULL on an already-open row is filled on a later parented turn, and a later different parent is ignored. Null means unknown (incapable harness, hostile or oversized value, or a conversation that names itself), never an error. List and detail session API responses expose parent_session_id with the same default. A well-formed parent id that has not been seen yet opens an empty parent row keyed the way the parent's own turns key ({principal}:{parent}), so a subagent that reaches the gateway before its parent's next turn still has a row to point at; that empty row is indistinguishable from a legitimately early parent until the parent itself arrives.
  • Published Vocabulary: For telemetry we align with OpenTelemetry GenAI's gen_ai.conversation.id, the only real cross-vendor standard for this concept. It is a telemetry attribute rather than a request field, so it is what we emit and document, while the wire-level intake is the precedence chain above.
  • Inactivity Closer (signal-less fallback): Without a native id, session identity would derive solely from the runtime principal, which for a durable managed-agent credential is machine-scoped and never changes, so every conversation on that machine appends forever to one row that is never ended_at. runtime_session_idle_timeout_minutes (default 720) bounds this: when the newest generation's last activity is older than the window, that row is closed at its own last activity (never at "now", so history is not rewritten) and the next request opens a new generation keyed <principal>:idle-<epoch>. It is strictly a safety net (a native session id always wins, so agents from levels 1-3 above are unaffected) and setting it to 0 disables it entirely. This is deliberately preferred over prompt-prefix inference, which was measured and rejected: two unrelated Codex sessions were 99.6% byte-identical (false merge) while one OpenCode session's consecutive requests shared zero messages (false split), i.e. it fails in both directions at once.
  • Operator Actions: Operators can end a session explicitly, which updates runtime state, emits audit and runtime-session events, and refreshes managed-agent summaries derived from the same principal.
  • Target Direction: Introduce a runtime-wide session abstraction that can represent flow executions, independent CLI/desktop agent sessions, and later enrolled workforce entities without making flow_execution the universal long-term session model.
  • Session search: Keyword chunks (session_search_document) are written on the same path as the source row. Optional embeddings are off until the account opts in; a bounded worker then fills vectors without sitting on the gateway request path. See docs/operations/session-embedding.md.

Gateway database ownership

HTTP gateway dependencies use the request Session only as an engine binding. Authentication creates, uses and closes its own Session in a database worker, returning frozen user/key/OAuth identity values. The bearer, password/key hashes and OAuth MCP credentials are not retained. Each subsequent database phase creates a fresh worker-owned Session; request dependency cleanup cannot close it.

Boundary Database work and values retained
Authentication and model listing Recheck bearer validity, runtime revocation, tenant scope and model bindings; close the worker Session before returning scalar results.
Preparation Resolve model permissions, runtime identity and existing budget estimates. Copy the selected model's configuration and budget decision into immutable values.
Policy evaluation Load current rules in a fresh short unit, then close it before detectors or approval waits. Reload rules at the initial stream pull and required final buffered-output checks.
Credential preparation Re-read the selected model through CRUD, resolve its current credentials, persist refresh/rotation, and close before inference. Provider callbacks receive only the required access credentials; Codex refresh tokens stay inside this phase.
Provider and stream OpenAI Chat Completions, Responses (including native passthrough and its transcode fallback), Codex, Anthropic and Gemini retain scalar model/auth/budget values. No database Session spans inference, stream pulls or retry backoff.
Completion and cancellation Record usage and budget rollups through CRUD in a fresh accounting Session. Close on success or failure. Unpriced-model live price lookup is scheduled with the model id only; the lookup worker re-reads the row through CRUD on its own Session so Alibaba overlay refresh can resolve stored credentials without the HTTP snapshot carrying secrets. Repeated cancellation drains a running worker before teardown; deferred recording retains the existing local once-only behavior.

OAuth credential rotation deliberately retains its serialized database lock across the bounded refresh HTTP call, because concurrent single-use refreshes would invalidate credentials. That lock ends before inference. Optional runtime summaries release the primary accounting unit before provider I/O, use their own credential Session and close it even on preparation failure. Already-loaded usage/runtime rows remain local to the same synchronous accounting worker; credential-bearing model rows are replaced with snapshots before summary I/O.

Credential replacement takes the same account-scoped secret row lock and reloads committed state before writing. Successful Claude/Codex rotations retain a bounded history of 64 consumed refresh-token HMAC-SHA-256 fingerprints in secret metadata, keyed with the instance secret and a purpose-specific prefix. Imports of those tokens are rejected, including unexpired local access bundles. Terminal provider failures are reused without resubmitting the rejected grant; transient failures remain retryable. preloop agents reconnect replaces credentials in place, clears terminal failure state, and consolidates that enrollment's legacy split subscription model rows onto one secret. It preserves enrollment identity and local configuration. This coordinates one Preloop instance; separate provider grants are required for independent refresh owners on other hosts or instances.

Internal replay, optimization and other caller-owned gateways keep their original Session and transaction boundaries. This ownership refactor adds neither atomic spend reservations nor provider-effect/retry deduplication.

Budget enforcement across editions

Basic BYOK spending policies run in the core gateway, including dedicated gateway processes. Enterprise extends the same enforcer for commercial controls and notifications; generic policies are evaluated once. A budget that cannot be evaluated is not a budget that was exceeded: when the price catalog has no entry for the requested model, no dollar comparison is possible, so a configured hard limit warns instead of rejecting. The request is served, its spend counts as zero, the usage row keeps pricing_available=false with real token counts, the caller gets X-Preloop-Warning: budget_pricing_unavailable: ... on both non-streaming and stream: true responses (the gateway resolves the model and runs budget preflight before the SSE headers go out), and an admin is paged on the existing unpriced-model alert path (24h per-model cooldown) so the catalog hole gets closed. A priced model over its hard limit is still refused. Subscription OAuth models with known zero marginal cost remain distinct from unpriced API usage. Requests with no applicable hard limit remain usable when pricing is unknown.

X-Preloop-Warning carries the non-fatal warnings a request produced, budget first, then alias collisions, joined with | and truncated to 256 characters. Streaming responses send their headers before the body generator resolves the model, so that branch does not carry the header.

X-Preloop-Usage-Id is the id of the ApiUsage row a non-streaming /openai/v1 request wrote, so a client can find the exact row the Cost page counts. It is absent when no row was written and on streaming responses, whose usage is recorded after the headers are sent. preloop models smoke prints it.

Checks estimate cost before dispatch; they do not reserve spend atomically. Concurrent calls can pass against the same remaining balance and exceed a limit when usage is recorded. These controls do not promise an exact concurrent ceiling.

Policy evaluation fetches account-scoped candidates once and resolves a managed agent or its owner only when a matching policy type requires that attribution. Legacy model-ID and model-alias policy rows remain enforced against the model's canonical spend bucket. User-scoped policies apply to spending attributed to agents owned by that user; they do not cover every directly authenticated call. An alias rename preserves legacy model-ID policy applicability to the current model bucket; it does not migrate historical spend from the former alias. New explicit alias-scoped policies must name an enabled gateway model available to the account. Existing account/agent/API-key alias policies remain readable if that alias is later renamed or disabled, but do not follow the model to its new alias. Review and explicitly replace those policies when changing model routing. Listing policies never rewrites aliases or historical spend.

Per-execution ceilings

BudgetPolicy bounds an account, flow, API key or agent over a period; it does not bound one run. A flow can additionally set agent_config.limits.{max_total_tokens,max_usd,max_turns}. The runtime API key minted for a flow run carries its flow_execution_id, and every gateway request is attributed to that execution, so before forwarding a request the gateway sums the run's own usage (api_usage.action_type='model_gateway', flow_execution_id match) and refuses with execution_budget_exceeded once a ceiling has been reached. Only already spent usage is compared, never a forecast: the request that crosses a ceiling completes once so the agent can emit its verdict, and the next one is refused. A run whose usage is entirely unpriced passes the USD check, for the same reason an unpriced model passes a hard limit. Refusal marks the execution FAILED with the budget_exceeded failure category and a message naming the ceiling. max_turns is counted at the gateway as one turn per model request, since no bundled runtime exposes a usable max-turns flag. Flows without limits are unchanged.

Historical repricing jobs

The billing repricing endpoint runs windows up to seven days in a worker thread. Larger windows create an account-scoped RepricingJob and publish its ID to the durable NATS JetStream queue. Acceptance requires a publish acknowledgement. The response includes job_id and status_url; clients poll GET /api/v1/billing/cost/reprice/{job_id} for queued, running, succeeded, or failed, instead of inferring completion from aggregate cost counts. The status endpoint requires manage_budgets and returns 404 across accounts.

The worker runs off the NATS event loop and renews message progress every minute. A database claim renews during scanning and execution rollup repair; after five minutes without progress, another delivery can recover the claim. Attempt numbers fence stale workers. A live duplicate remains unacknowledged, and terminal deliveries are idempotent. Three abandoned attempts terminate with an error. Caught failures are terminal and require a new submission. Some batches may already have committed when a job fails or is recovered; reported counters belong to the latest attempt, and stalled flags queued or running jobs whose progress is overdue. Job records currently have no automatic retention pruning.

Repricing uses currently active account overrides retroactively, preserves provider, reconciled, imported, and subscription costs, and never rewrites budget spending. Explicitly unavailable pricing metadata is eligible even when a legacy row already contains zero cost. Both bulk and single-row repair update override provenance and pricing availability when cost and source are unchanged. Existing request-time budget decisions and limits are preserved.

Unresolved rows backed by a trusted OpenRouter model also attempt a bounded read of https://openrouter.ai/api/v1/generation using the stored upstream_request_id and account-scoped API credentials. The pass makes at most 50 calls with a three-second request timeout and a 30-second cumulative lookup budget. It memoizes generation results, suppresses failed credentials, and stops requests after a rate limit. Recovered per-request costs are tagged provider with retrieval provenance; explicit zero is valid. Missing IDs, unavailable records, ambiguous costs, and exhausted lookup budgets remain unpriced. Dry runs can preview these provider reads without changing usage. The provider_lookup response counters explain recovery and unresolved work.