Cost Analytics and Budgeting¶
Editions: OSS. Contributor documentation for this repository.
Cost analytics turns gateway telemetry into explainable spend and budget health. This chapter covers the ApiUsage ledger, OSS API/UX boundaries, and the Enterprise plugin split.
Progressive reporting¶
The Cost console requests GET /api/v1/cost/summary?include_breakdown=false
for its first paint and previous-period comparison. This keeps account totals,
budget, pricing/unpriced context, and separately reported imported totals, but
skips the grouped breakdown queries. Settings and model metadata load separately
and do not block those totals.
Callers can select details with repeated breakdown parameters: models,
flows, sessions, tools, days, or imported. For example,
?breakdown=sessions&breakdown=flows loads the Agents tab without computing
tool costs or the daily timeseries. With no new parameters, the endpoint
retains its full historical response. include_breakdown=false takes precedence
over a selection. Unselected arrays are empty because they were not requested;
clients must distinguish this from a loaded section with no data.
Agents, Sessions and Users share an in-flight/session breakdown within the current view. Tools and user ownership load when their tabs are opened. Imported details load separately when imported totals identify visible content. Each section has its own loading/error/retry state. Range changes invalidate loaded sections and reject late responses, then reload the tab that remains selected. All details use the effective period returned with the initial totals. The console does not persist previous-period results across account sessions.
This changes request scheduling and selected query execution only. Account isolation, history policies, ledger accounting, attribution, reporting limits, and full-query ordering remain unchanged. It adds no rollups or response cache.
Query shape¶
Session breakdowns aggregate raw api_usage rows by session and model first.
Session name, agent, flow, and principal labels are joined onto that aggregate.
The response limit applies after the full aggregation, so a capped session list
still carries complete totals for each returned group. Daily series aggregate
in a materialized day bucket, then sort those buckets. Per-user windows filter
runtime_principal_id through ix_api_usage_account_principal_id_ts
(account_id, runtime_principal_id, timestamp for model_gateway rows).
ix_api_usage_account_principal_ts still leads with principal type. Accounting
rules are unchanged: replay-validation rows stay excluded, retries stay included
unless the caller sets exclude_retries, and there is no daily rollup or
response cache.
Cost Analytics and Budgeting¶
- Purpose: Turn model usage telemetry into explainable spend, enforceable budgets, and optimization guidance.
- Canonical Ledger:
ApiUsageremains the source of truth for model call tokens, estimated cost, provider, model, runtime principal, API key, flow, managed agent, and runtime-session attribution. - Idle Cache Expiry:
preloop.services.context_analysisextendsCacheProfilewithCacheIdleExpiryEventrows when consecutive content-stable gateway calls are separated by more than the provider idle TTL (Anthropic 5m, OpenAI 10m, Gemini 1h, DeepSeek 2h) and ApiUsage shows a cache_read collapse plus cache_creation spike. Extra cost is(write_price_per_1k - read_price_per_1k) * rewritten_tokensfrom the vendored catalog; optimize/replay surfaces only measured, per-session figures. - Accounting Self-Check:
GET /api/v1/cost/healthverifies the accounting chain end-to-end per account over a lookback window (gateway traffic seen → streaming requests record tokens → costs priced → provider-reported usage share → audit events present), so silent accounting breakage (like streaming rows recording 0 tokens) is caught immediately instead of weeks later. - Effective Price Read-Back:
GET /api/v1/ai-models/{id}/pricing(view_ai_models) answers what one model is priced at right now and where that number came from, resolving in the gateway's own order: an account price override, then the pricing configured on the model, then the vendored catalog, thensource="none".POST /api/v1/ai-models/{id}/pricing/fetch(edit_ai_models) reads a provider's published price (OpenRouter's public model list, or Alibaba Cloud Model Studio's native catalog on USD sites) and never writes it: a fetched number is confirmed by a person through the price override endpoints before it changes what spend means. Both are account-scoped through the model, and a malformed stored price reads as unpriced rather than failing the page. - Priced, Unpriced, and Zero: each
usage_by_modelrow carriesunpriced_request_count(tokens spent with no price at all, so that cost is missing from every total),zero_priced_request_count(a price was applied and it was exactly zero, so nothing is missing),failed_request_count, andlast_request_at. The split exists because a $0.00 total means two different things, and only one of them is an accounting hole.unpriced_request_countreuses theget_gateway_usage_summarycondition, so the per-model counts sum to the account total. - OSS API Surface: Core endpoints should provide aggregate summaries, grouped breakdowns, raw usage drill-downs, and budget-health alerts derived from gateway account/flow limits. Core endpoints also provide runtime-session optimization recommendations, one-click apply, and replay verification (
preloop/api/endpoints/session_optimization.py), with hosted-model analysis gateable via thepreloop.services.optimization_gatingauthorizer hook. Enterprise billing plugin endpoints provide budget policy CRUD, enforcement, and model price override CRUD behind feature flags. - OSS UX Boundary: Open source should answer "how much was spent?", "who or what spent it?", and "which budget applies?" with enough drill-down to inspect the related session timeline.
- Enterprise UX Boundary: Enterprise should answer "why was it spent?", "was it worth it?", and "how could it be optimized?" at scale with LLM-assisted reviews, anomaly detection, forecasting, showback/chargeback, credits/promotions, exports, and workflow automation.
- Default AI Model Use: Enterprise session-value analysis should call the account's default AI model through the Preloop Gateway, producing an auditable meta-usage record for the evaluation itself. The analysis should reference redacted session summaries, gateway events, tool calls, approvals, and final outcomes rather than unrestricted raw prompts.
- Plugin Boundary: Backend features beyond OSS summaries and budget-health tracking must live in Enterprise plugins under
./plugins/, likely extendingplugins/billing/for budget policy enforcement, pricing overrides, FinOps, credits, promotions, forecasting, exports, and value-review jobs. The shared frontend should gate those panels with feature flags. - Budget Actions: Core enforcement should continue to block or warn before upstream dispatch. Enterprise plugins can add escalations, Slack/mobile notifications, approval requirements for expensive calls, and post-hoc anomaly workflows.
Cost and cycle time per tracker issue¶
preloop.services.issue_cost_rollup rolls execution cost, tokens and pull
request cycle times up to the tracker issue, across flows. The per-flow Cost
page is unchanged.
- Tables:
issue_cost_rollup(one row per account, tracker and issue key),issue_cost_execution(one fact per execution id, so a replay or a rebuild upserts instead of double counting) andissue_cost_pull_request(publication, approval and merge times, the claiming issue and an ambiguity flag). Row sums are always recomputed from the facts. - Attribution: first match wins: issue lifecycle, resume lineage, delegated parent, retry parent, an issue trigger subject, a pull request already claimed by one issue, exactly one closing reference. Anything else is unassigned. A pull request claimed by two issues is marked ambiguous, and executions linked only through it move to the unassigned bucket.
- Write hooks: the orchestrator terminal hook, the execution monitor's
stale pass and the crashed local dispatch path (terminal statuses written
outside the orchestrator),
record_opened_pr(publication time),process_webhook_event(approval and merge times, never creating rows) andsync_execution_cost_rollup(repricing). Each runs in a savepoint and never fails its caller. - PR opened time:
issue_cost_pull_request.opened_at_sourcesays whereopened_atcame from.forgeis the pull request's owncreated_at, read through the tracker'slist_open_pull_requests_by_source_branchon the branch lookup bind path (GitHub, GitLab and Bitbucket) or from any later pull request webhook; it replaces a Preloop time even when that is earlier.bindis the time Preloop bound the pull request to the run andrun_endthe end of the publishing run; both only fill an empty value. The issue row carries the source aspr_opened_at_source. - Rebuild:
POST /api/v1/cost/by-issue/rebuildrecords finished executions of a window of at most 92 days that have no fact yet, each in its own savepoint. It is the recovery path for executions that ended outside the orchestrator or whose hook failed. - Scheduled rebuild: the API role runs the same rebuild every
ISSUE_COST_REBUILD_INTERVAL_SECONDS(default 3600) for executions that started in the lastISSUE_COST_REBUILD_LOOKBACK_HOURS(default 72), at mostISSUE_COST_REBUILD_MAX_EXECUTIONS_PER_ACCOUNT(default 500) per account per pass. Each account is rebuilt in its own transaction under apg_try_advisory_xact_lock, so replicas skip an account another one is rebuilding. The pass then re-reads the estimate of recently active issues from their synced issue rows.ISSUE_COST_REBUILD_ENABLED=falseturns it off. Older history still needs the rebuild endpoint. - Estimate: the human estimate as the tracker states it, never
derived, for comparing AI cost with the estimate. Hours come from Jira
Original Estimate (
timeoriginalestimate) or GitLabtime_estimate. Points, and hours on trackers without a native field, come from the tracker'smeta_data.issue_estimateconfiguration:points_field(an issue field such as a Jira story points custom field or GitLabweight),hours_label_prefixandpoints_label_prefix(labels such asestimate:4horsp:3; two labels with different values are no estimate). Set it withPUT /api/v1/trackers/{id};meta_datais replaced as a whole, so send the existing keys too. Values are read from the trigger payload when it is about the issue and from the synced issue row (Jira and GitLab store the native fields inmeta_data.estimate_fields), the synced row winning. A reading that states nothing never clears a stored estimate. Empty when the tracker has no estimate. - Cycle-time definitions: the report measures from the first agent run
to the merge, never from ticket creation.
first_event_atis the earliest start of any execution attributed to the issue, across all flows. It is not the ticket's creation time: time a ticket waited before its first run is not measured.pr_opened_atis when the linked pull request was opened, with the provenance inpr_opened_at_source(forge, elsebind, elserun_end; see below).approved_atis the earliest recorded approval event (GitHub review with state approved, GitLabmerge_request_approved, Bitbucket Cloudpullrequest:approved) at the time the forge reported for it. It is not verified mergeability: required checks, approval counts and branch restrictions are not evaluated, so a pull request can be approved and still not mergeable.merged_atis the earliest recorded merge event (Bitbucket Cloudpullrequest:fulfilleduses the pull request'supdated_on).first_event_to_pr_opened_hours,pr_opened_to_approved_hoursandapproved_to_merged_hoursare elapsed UTC hours between those milestones, rounded to two decimals, blank when a milestone is missing or the two are out of order. No interval runs from ticket creation, and approval-to-merge is not a time-to-mergeable figure. Ticket creation to an observed ready-for-merge state is a separate, explicitly scoped metric.- Redelivered or out-of-order webhooks never move a milestone: the earliest approval and merge times win, and approval and merge events are never counted as executions.
- Report:
GET /api/v1/cost/by-issuefilters issues by first event time (start_dateinclusive,end_dateexclusive) and shows their lifetime totals. An issue whose first run started before the period is excluded even when its merge falls inside it. A flow filter restricts cost, tokens, runs and execution ids to that flow's executions; the milestones and intervals stay those of the whole issue. The per-project and per-flow summaries are sums of the rows./unassigned/executionslists the runs in the unassigned bucket for the same filter./exportreturns CSV (issue grain plus one unassigned row) or JSON (with execution ids). - Cost coverage:
estimated_costis the subtotal of the runs that carry a cost, so on its own it cannot tell a free ticket from an unpriced one. Every issue row, every project and flow summary and the unassigned bucket therefore also reportcost_coverage,known_cost_run_count,unknown_cost_run_countandattributed_cost_usd. Coverage iscompletewhen every contributing run has a cost,partialwhen both kinds are present andunknownwhen none has one; an empty bucket isunknownwith both counts zero, and a known zero counts as known.attributed_cost_usdis the subtotal only forcompletecoverage and null otherwise, so a partial or unknown bucket is never read as a total. The counts come from the same account-scoped fact aggregates as the sums, so they follow the report's filters (a flow filter prices only that flow's runs) without a per-execution usage query. - What coverage is not: it describes execution-cost availability, never
invoice accuracy. A
completerow is still an estimate priced from published model rates. The three other cost signals stay deliberately outside these numbers: the premium-request counts a host CLI run reports (host_exec_usage, shown on the execution), the daily GitHub Copilot import (copilot_usage_import, account level and never attributed to a ticket) and any per-seat subscription price. No daily import dollars are added to an issue total and no seat charge is inferred per ticket. Existing facts keep their stored cost: an explicit historical zero stays known unless its producer is independently shown to be wrong. -
Export columns: CSV columns, in order:
tracker,issue_key,title,project,estimated_cost,total_tokens,run_count,failed_run_count,first_event_at,pr_opened_at,approved_at,merged_at,first_event_to_pr_opened_hours,pr_opened_to_approved_hours,approved_to_merged_hours,issue_url,pr_url,pr_opened_at_source,estimate_hours,estimate_hours_source,estimate_points,estimate_points_source, then the appendedcost_coverage,known_cost_run_count,unknown_cost_run_count,attributed_cost_usd. Blank means unknown, never zero. The JSON export'sissues[]objects carry the same fields (null for unknown) plusexecution_ids. Estimate sources arejira:timeoriginalestimate,gitlab:time_estimate,<tracker type>:<points_field>orlabel:<prefix>.The four coverage columns are appended, so an older consumer keeps reading the same names in the same order and
estimated_costkeeps its type. What changes is the interpretation, not the shape: a consumer that adds the issue rows together now sees the priced subtotal only, which understates a ticket whose runs were subscription-backed.tracker,issue_key,title,project,estimated_cost,...,cost_coverage,known_cost_run_count,unknown_cost_run_count,attributed_cost_usd GitHub,example-org/example-repo#12,Add the export button,Example,2.0,...,partial,1,1, GitHub,example-org/example-repo#13,Subscription-backed work,Example,0.0,...,unknown,0,2,The second row is not a free ticket: two runs carry no per-run price, so the row reports no attributable total. A consumer that wants a total only where one exists sums
attributed_cost_usdand treats the blanks as unknown. * Worked example (synthetic): a Jira ticketREC-7created at 08:00 UTC, a Bitbucket Cloud pull request and five runs, as recorded by the reconciliation fixturebackend/tests/issue_cost_reconciliation.py:Time (UTC) Event Cost 08:00 Ticket created (not measured) 09:00 Implementation run starts and fails 1.00 09:20 Distinct retry publishes the pull request 0.25 10:00 Bitbucket created_on(Preloop bound it at 10:05)10:30 Review run, triggered by pullrequest:created0.50 11:00 Repair turn resuming the review 0.10 11:30 Seat-backed review run, no per-run price null 12:00 pullrequest:approved(delivered twice, after the merge)13:00 pullrequest:fulfilled(delivered twice)The issue row reports
first_event_at09:00,pr_opened_at10:00 with sourceforge, intervals 1, 2 and 1 hours,estimated_cost1.85 from four priced runs,run_count5,failed_run_count1,cost_coveragepartialand a nullattributed_cost_usd. A missing Jira original estimate stays blank, and a story points field set to 0 reports 0. The 08:00 to 12:00 span is never reported: it would claim a mergeable time the report does not observe. * Reconciliation procedure: for one issue, export JSON and CSV with the same filter. Check that the JSONexecution_idsare exactly the runs you expect (one per unique execution; a retry is its own run), thatestimated_costequals the sum of their known costs and the run counts match, that the per-flow and per-project summaries add up to the rows, and that the unassigned bucket holds the runs that could not be tied to one issue. Compare the four timestamps (UTC offsets included), the three intervals and the blank or null cells between the two exports. A run whoseestimated_costis null is counted inunknown_cost_run_count, not as zero.
Spend outlier alerts¶
preloop.services.spend_outliers flags a developer or session whose spend
departs from the usual pattern. It reads gateway ApiUsage rows
(action_type='model_gateway', replay validation excluded) grouped by user,
UTC day and model. It does not add flow_execution.estimated_cost, because
those calls are already usage rows.
- Daily spend: spend on UTC day D is at least
daily_multiple(default 3) times the median of the days with spend among the previous 28. The rule needsmin_history_days(default 7) such days, and a zero median never fires. - Model mix: one model matching a
top_tier_model_prefixesentry (case insensitive, with or without aprovider/prefix) is more thantop_tier_share(default 0.5) of the developer's spend on both D and D-1. - Session: one runtime session costs more than
session_cost_threshold_usd. The rule is off while that is null.
The daily rules run at 00:30 UTC for the day that just ended. The session rule
runs every 15 minutes over sessions active in the last two hours. Settings
live in spend_outlier_settings, one row per account, and are edited under
/api/v1/attention/spend-outliers/settings (manage_budgets to write,
view_cost to read).
Fires once. Each finding is a row in spend_outlier_finding, unique on
(account_id, fingerprint) and written with ON CONFLICT DO NOTHING, so a
rerun, a retry or two workers racing record it once. The attention item id is
stable per rule and developer (spend:<rule>:<user_id>, or
spend:session_cost:<session_id>). The fingerprint names the UTC day
(<rule>|<user_id>|<YYYY-MM-DD>, or session_cost|<user_id>|<session_id>).
Dismissal. Cards use the existing attention dismissals. A dismissal hides
the card while its fingerprint matches, so a developer who is still an outlier
on the next day gets a new card. A snooze is the exception: for spend cards an
unexpired snooze hides the card whatever the fingerprint, until the snooze
ends. The dismissal endpoints stamp dismissed_at on the matching finding,
and a restore clears it.
Digest. build_spend_outlier_digest_section(db, account_id, now,
start=..., end=...) returns the findings detected in one half-open
[start, end) window, one entry per fingerprint, each marked dismissed when
a dismissal or a snooze covers it. The two are read at different moments: a
dismissal is read as it stands, so a finding dismissed after the window closed
is still reported as dismissed, while a snooze is resolved as of the window
end, so one that only runs out afterwards still covers the window. Without
start and end the window is the seven days ending at now; with them it
is exactly that window, and both bounds are always applied, so a finding
detected at or after the end of the window is not in the section. The window is
reported back as window_start and window_end. Display names, session titles
and dismissals are resolved within the account, so a row that points at another
account's user or session shows no label rather than that account's. Every
entry keeps the numbers its rule recorded, including imported dollars;
overlapping findings are never summed into an account total. It is the section
for the weekly digest service, which is resolved through the plugin registry
and lives outside this repository.
Imported spend. Spend that does not pass through the gateway enters
through register_imported_spend_source. Cards and digest entries that
include such dollars say they are not metered by the gateway. The one
production source is preloop.services.copilot_spend_source, registered by
the evaluate_spend_outliers task before each pass (idempotent, no HTTP
router imported, no GitHub call). It reads stored Copilot premium-request
rows for the connection's current organization, keeps daily per-user rows in
USD whose login has a row in copilot_user_mapping (operator-written, one
active same-account user per login, several logins per user allowed), nets
signed amounts per user, day and model, and returns the positive nets as
source='copilot'. Seat rows, usage metrics, organization totals, the
unattributed residual, unmapped logins and rows without an amount are left
out and reported as counts by GET /api/v1/cost/copilot/spend-coverage.
Nothing is written back to api_usage, budgets or issue rollups.
Replay and supersession. Imported days arrive three days late and can be
corrected, so for accounts with an active Copilot connection the daily pass
evaluates the 28 most recent completed UTC days (REPLAY_WINDOW_DAYS) rather
than yesterday alone. evaluate_days loads the span once, judges each day
against its own 28-day history, and reconciles the stored findings for
exactly the two daily rules and those days: an unchanged finding is left
alone (dismissals and snoozes keep applying), changed evidence is written
onto the existing row with detected_at kept, and a day that no longer
qualifies gets superseded_at and superseded_reason and is filtered out of
the open list and the digest while its row remains. superseded_at and
dismissed_at are independent. When an imported source raises, nothing is
reconciled and only yesterday is judged from gateway spend; the pass reports
the account as incomplete.
Reviewed price publication¶
After an initial rollout and explicit configuration, each API, dedicated gateway, and worker polls the same trusted HTTPS price artifact. A reviewed publication updates supported flat token rates, native DeepSeek UTC peak/off-peak tariff revisions, and dedicated Alibaba USD regional token tiers without deploying application code. Unknown policy structures require an estimator change, boundary tests, and a deployment. Refresh validates evidence, effective dates, model scope, and historical tariff continuity before replacing the current map; existing usage records, account overrides, and provider-reported costs are unchanged. On failure it retains last-known rates as potentially stale estimates and logs the failure. The public weekly model-price review preset and idempotent installer bind the account's existing model and repository, prepare an evidence-backed PR and regional feed, and report providers or cache policies that could not be verified. Alibaba prices land in the dedicated region store of each process; scoped regional allowlists can admit newly reviewed SKUs, while freshness and effective dates prevent stale native overlays or newer tariffs from corrupting historical estimates. See configuration and publication.