Durable implementation feedback¶
Editions: OSS, Cloud, Enterprise. Unless stated otherwise, everything on this page ships in OSS.
The Automated Issue Implementation preset can keep a PR moving through review and CI without leaving an agent container waiting. Each repair gets a new FlowExecution, its own execution budgets and fresh credentials. The implementation thread keeps the PR branch and, when recovery files are available, the exact native conversation across those turns. Reviewers remain separate flows and conversations. Merging remains a human action.
Enable a subscription¶
New copies of preset 011 enable agent_config.feedback. Existing saved flows
keep their configuration until the operator opts in or deliberately applies a
preset update. The issue's agent-ready label is an intake condition; feedback
is routed by the recorded account, flow, tracker, numeric provider repository
identity and PR number, without requiring that label on the PR.
agent_config:
feedback:
enabled: true
debounce_seconds: 30
max_turns: 5
max_cost: 100
max_no_progress: 2
max_age_hours: 168
ci_deadline_seconds: 3600
# Reconciliations to wait out a provider infrastructure failure on one head.
max_ci_infra_retries: 3
repair_early: false
# Usernames or GitHub/GitLab app slugs. ``preloop`` matches preloop[bot].
# Numeric actor ids still work. Comment markers never grant trust.
trusted_reviewer_ids: ["preloop"]
implementer_actor_ids: [67890]
# Optional project policy; provider rules can add required gates.
required_checks: ["Backend Tests", "UI Tests"]
required_approvals: 1
In the console, edit a flow and enable Continue implementation after PR review
or CI failure under PR review and CI follow-up. Trusted reviewers start as
preloop, which matches reviews posted by the Preloop GitHub App
(preloop[bot]). A different app, such as preloop-staging, has to be listed
by its slug. Numeric actor IDs still work. Enter the implementer's numeric
GitHub or GitLab actor IDs, then choose limits for repair
turns, cumulative estimated cost in USD, lifetime, and feedback debounce. The
initial values match the policy defaults above. Saving an existing flow without
opting in leaves follow-up disabled. Other policy fields set through the API are
preserved when editing these controls. A compatible saved checkpoint is required
for native conversation resume. An existing publication without that state
requires explicit permission to start a fresh conversation. Turning on the option does not merge a PR.
For an existing PR, open the finished execution that published it and choose Set up PR follow-up. Failed, timed-out and cancelled executions can also be selected when they produced a PR before reporting stopped. If the execution has no recorded publication, enter the existing PR URL and its source branch. The server reads that exact PR through the execution's original tracker and requires both sides of the PR to belong to its original repository. It checks the source branch and current head; a recorded PR association cannot be replaced with another one. No branch search or new PR creation is performed.
This fetches a read-only preview of the exact PR, branch, current head and recovery options. Confirm a fresh conversation explicitly if its saved state cannot be resumed. If the head changes, review the refreshed preview and confirm again. The action requires enabled flow feedback and deployment support for saved execution state uploads; unavailable prerequisites are shown in the preview. A failed network response triggers a status refresh, not an automatic second adoption request.
Enabling feedback applies to future executions. The server records opt-in when an execution starts; turning the setting on does not discover and repair an old PR backlog. An older publication requires explicit adoption. Turning feedback off pauses future repairs without cancelling a running turn. Re-enabling keeps consumed turns, cost, no-progress history and the deadline. Current reviewer trust and budget settings apply on every reconciliation and are checked again when reserving a repair. Reducing the maximum age can shorten the original deadline; increasing it never extends an existing subscription.
Upgrade an older issue-only flow¶
An older saved implementation flow may listen only to issue_labeled. Updating
the installation or retrying that execution does not add review triggers. To
continue its PR:
- Edit the saved flow, enable PR review and CI follow-up, and add the review
app's slug (for example
preloop) to Trusted reviewers. This subscription handles review and CI events independently of the flow's issue trigger types. - Keep
agent-readyas the intake filter; do not add that label to the PR just to make review feedback work. - Open the original finished execution and preview PR follow-up. If reporting was lost, supply its existing PR URL and source branch. Check the current head and explicitly acknowledge a fresh conversation when recovery files are unavailable, then adopt that publication.
A fresh conversation uses the published branch and original issue criteria. It does not restore missing native session files or change the failed execution to successful. An older execution without a subscription needs this explicit adoption; turning feedback on alone does not restart it. Read-only preview and a failed adoption do not change the PR association or create a subscription.
The first repair of a publication whose publisher stored no native session also
continues on that published branch, including when checkpoint uploads are
disabled. Later repairs also use the published branch when no recoverable native
checkpoint was stored. An expired, mismatched or incorrectly bound stored
checkpoint still fails closed. A deployment that
replaces the chart's default worker pool must subscribe one pool to
reconcile_flow_feedback; otherwise the scheduler publishes reviews that no
worker reads.
Use the reviewer app's slug or numeric actor ID in trusted_reviewer_ids.
preloop trusts preloop and preloop[bot] only. Unlisted bots and the
configured implementer actor are ignored. A copied HTML
review marker never grants trust. All comment and CI text is untrusted task data.
Cost is cumulative estimated execution cost in USD, with existing execution
budgets enforced independently. No-progress detection compares the PR head
before and after a repair. The execution result's continuation object shows
thread state, consumed turns/cost, pending feedback, head and stop reason.
The execution detail page marks each repair as a Continuation and links to the original publishing execution. Both pages provide issue and pull request links, plus an ordered list of follow-up executions with their status and start time. The list shows up to 100 continuations and links to the flow's paginated execution list when there are more. Delegated executions remain a separate tree.
Publication registers an internal subscription using existing repository
webhooks. No webhook is installed for an individual PR. Registration recovery
and periodic provider reconciliation cover feedback that races with publication
or whose webhook was lost. The sync scheduler publishes reconcile_flow_feedback
every 15 seconds; workers reconcile bounded batches. Default feedback debounce
is 30 seconds. Duplicate deliveries and check/workflow notifications do not
create duplicate execution turns.
A repair that fails before any agent work (the resume was refused or the launch
errored: no session, no runtime and no tokens) does not consume a repair turn.
Its feedback stays pending and the next attempt waits out a growing backoff
(4 minutes, then 16, up to 6 hours) instead of retrying on every
reconciliation. While waiting, the continuation shows
stop_reason: resume_launch_retry. A repair that reached the model and then
failed, for example out of memory, still counts as a turn.
A PostgreSQL row lease protects each thread. Creating the next PENDING execution and assigning its feedback receipts is one transaction. If dispatch fails or a worker crashes, normal execution recovery dispatches the same execution ID. Feedback that arrives during execution stays pending. A stopped, cancelled or aborted publisher or repair stops its subscription; a cancelled execution with an older publication can still be adopted explicitly. The agent runner exits between turns; CI waiting and stuck-job deadlines belong to the scheduler.
Provider gates¶
GitHub reconciliation reads the current PR head, checks, legacy commit statuses,
submitted reviews, unresolved inline review threads and conversation comments.
It combines configured required checks with branch protection and effective
ruleset check/review requirements; flow configuration cannot lower repository
requirements. GitHub's explicit 404 Branch not protected response establishes
empty classic protection; an ambiguous 404 does not establish that requirements
are absent. Nonempty submitted review summaries in the COMMENTED state enter
feedback without replacing the reviewer's previous approval or changes-requested
verdict. A failing check run contributes its own
bounded, redacted output title/summary/text as diagnostic evidence. For failing
GitHub Actions checks, reconciliation reads at most two job-detail records on
the bound repository. The job must match the current head and check-run ID.
A failure in the provider-owned first setup step is infrastructure; a failing
user step is code evidence. Missing, stale, denied or over-budget job evidence
blocks classification instead of starting a speculative repair. Job URLs never
become outbound request destinations. startup_failure uses the existing
bounded infrastructure retry/escalation policy.
GitLab reconciliation reads MR notes, current-head commit statuses, approvals and blocking discussion state. It also reads the current head's pipeline (from the MR head pipeline, otherwise the newest pipeline for that exact SHA) and that pipeline's jobs. Retried attempts, jobs from another SHA and jobs from another project are discarded, and the newest attempt of each job name decides. For at most three failing jobs, a bounded redacted tail of the job trace is read as diagnostic evidence. The trace is streamed and discarded as it arrives, so an enormous log never enters memory whole, and credentials are redacted before the tail is cut. A missing or forbidden trace is simply absent. When the pipeline has no readable job (a configuration error, or jobs the token cannot list), its own status and any explicit provider failure reason are the evidence instead. Job and pipeline reads stay inside one provider page, like notes and statuses.
Both paths recheck the head and open/closed state after reading gates and stop
repairing closed or merged PRs, including closure during the gate reads. GitLab
approval readiness requires both the provider approval rules and any configured
required_approvals minimum, counted by distinct approving users. GitHub legacy
commit-status pagination blocks readiness just like truncated check-run results.
Only current-head check failures trigger repairs, and only when provider details attribute the failure to the branch. Pending or missing required checks wait until the CI deadline, then report an explicit blocked state. Current cancellations, startup failures, permission requirements and unknown outcomes block readiness instead of inviting speculative code edits. Neutral, skipped and allowed-failure outcomes follow provider semantics. A failed required check never becomes ready simply because a webhook was missing.
Terminal failure classification¶
Every failing required check is classified from provider metadata, never from the log text (a trace is untrusted task data and cannot request a repair):
| Evidence | Outcome |
|---|---|
GitLab job script_failure/test_failure, GitHub Actions failing user step, or a non-Actions check-run failure |
code failure: one coalesced repair round |
| Runner, API, scheduler, image-pull and similar platform reasons | infrastructure: bounded wait, then ci_infrastructure_failure |
GitLab timeout reasons, GitHub timed_out without a more specific failed-step reason |
infrastructure: bounded wait, then ci_timeout |
| Quota, archived project, blocked user, protected environment, upstream permission reasons | ci_permission_required, a human must act |
unknown_failure, an unrecognised reason, or a failing check with no readable job |
ci_failure_unclassified |
| A retried attempt whose newer attempt decided the check | ignored |
Infrastructure failures never consume a repair turn. They are retried for
max_ci_infra_retries reconciliations of the same head (default 3, reported as
ci_infrastructure_failure_retry/ci_timeout_retry), then the thread blocks with
the reason above. A new head or a recovered rerun clears that allowance. Review
feedback that arrives while CI infrastructure is broken still repairs normally.
Flows without durable feedback still use the legacy webhook resume path. That
path ignores startup_failure, timed_out, and action_required rather than
starting a code repair without job evidence. Its ordinary failure handling
is unchanged; bounded job-detail enrichment applies to durable subscriptions.
Readiness requires passing checks and review gates on the current head. Provider permission errors, pagination beyond the bounded reconciliation window, and ruleset workflow/code-scanning gates that require additional evidence prevent readiness with a reason. Explicit permission denials on GitHub branch protection or ruleset discovery still allow repairs to fully read, current-head review and CI feedback. They never establish that repository requirements passed. Other provider failures, rate limits or incomplete feedback block repairs as well. Resolve the blocker or provide the required gate integration; Preloop does not interpret unavailable evidence as approval. The PR remains open for manual merge.
Native conversation checkpoints¶
Native session data is separate from workspace recovery. A versioned manifest names the harness/version, explicit session ID, implementation thread, file hashes, size limit and expiry. Codex checkpoints contain the selected rollout and verified child rollouts. OpenCode 1.2.6 and 1.18.29 use SQLite: a consistent read transaction exports only that session graph into a new database, without unrelated sessions, account credentials, share secrets or persisted permission grants. A database copy followed by deletion is insufficient because unused SQLite pages may retain foreign data.
The native artifact uses the encrypted artifact service and scoped direct HTTP capabilities. It is never sent through logs, generic trigger JSON, MCP tools or a shared account volume. The resolver authorizes account, flow, thread, reserved current execution and latest prior execution before issuing restore capability. Providing another issue's execution ID or artifact reference does not authorize its session. Private runners must advertise native checkpoint support; a runner without it reports a cold handoff and does not upload its home directory.
Restore begins in an empty session directory. Missing or expired recovery files
stop a durable native resume. When no native artifact was stored or checkpoint
uploads are disabled, the controller can authorize a fresh conversation on the
bound published branch and report cold_handoff. An explicit adoption applies
only to its selected original publication. Later native resumes need their own
workspace and native checkpoints. Corrupt, mismatched,
unsupported or incompatible existing state produces resume_failed. A failed
native CLI resume preserves the checkpoint and does not silently select another
conversation. The execution result records native_resume, cold_handoff or
resume_failed. Native manifests default to seven days; the artifact service's
native retention must cover that window independently of workspace retention.
Codex's CLI is pinned to npm release 0.153.4, tested against the shipped universal
image. OpenCode is pinned to 1.18.29, with agent_config.opencode_cli_version
as its exact-version override. Its current SQLite event/context tables are scoped
to the selected conversation; account, credential and share tables stay empty.
agent_config.codex_cli_version accepts an exact release version for an
intentional upgrade. Upgrade tests must repeat the two-turn image smoke. A
checkpoint from another CLI version is rejected explicitly. OpenCode's image
version and storage schema are validated through the native manifest. Affinity
and completion reminders always use explicit IDs, never latest-session flags.
Wrappers install the selected CLI before one native restore, enter the primary
checkout after setup, and log the actual CLI version and configured image reference.
A configured image tag is not proof of the resolved runtime image digest.
Deployment prerequisites¶
Enable checkpoint uploads only after deploying the transaction-resilience backend
and its EE companion. In particular, the artifact quota lock must use
FOR NO KEY UPDATE so checkpoint inserts do not recreate account foreign-key
contention. Apply database migrations before starting the updated API and flow
workers. Merely enabling feedback on a saved flow does not enable artifact upload.
The chart already supports shared extraEnv on the API, gateway, execution workers
and scheduler. Its defaults leave FLOW_ARTIFACT_DIRECT_UPLOAD disabled. The
optional native checkpoint overlay
enables it and retains both workspace and native artifacts for seven days. Copy
and review that file with your installation values; do not enable it merely by
setting an environment variable on the Helm client or CI job.
The example overlay caps compressed uploads at 64 MiB through
WORKSPACE_SNAPSHOT_MAX_BYTES, which applies to both artifact kinds, and sets
gateway.proxy.bodySize: "80m" for the ingress and console proxy. This avoids
sending workspace archives through the legacy 2 MiB Kubernetes log channel.
Measure a representative workspace and native-session archive locally before
choosing a different limit: repository history and generated assets can exceed
it. Check for explicit ingress annotation overrides in the installation values.
Oversized archives
fail explicitly; increase application and every proxy limit together only after
checking memory and database capacity. Expanded archives retain their separate
FLOW_ARTIFACT_EXPANDED_MAX_BYTES limit (default 2 GiB), and
FLOW_ARTIFACT_ACCOUNT_QUOTA_BYTES limits stored account data (default 4 GiB).
Retention consumes that quota, so cleanup and database headroom must cover the
selected retention window. The checkpoint interval defaults to 300 seconds.
Upload buffering, encryption and database driver copies can require several times
the compressed archive size in memory. The expanded-size limit is validated with
streaming reads, but it still bounds restore disk usage and processing work.
Keep upload concurrency bounded during initial validation; increasing a size
limit alone does not establish sufficient memory or storage headroom.
Helm merges maps but replaces lists. Merge these entries with any existing
extraEnv, preserving database, private-CA and other installation entries. An
additional values file must not silently replace that list. Render the combined
values and verify that the same direct-upload flag reaches the API and execution
workers. Confirm the existing signing/encryption Secret references remain intact;
never print key values in CI logs.
All API and worker replicas must share a stable SECRET_KEY, used to sign and
verify scoped upload/download capabilities. Artifacts use the existing encryption
configuration: a stable SECURITY__ENCRYPTION_KEY when configured, otherwise the
existing encryption key derived from SECRET_KEY. Preserve the installation's
current mode and keys. Switching to a new dedicated key or rotating either key is
not part of enabling checkpoints and can make existing encrypted data unreadable.
If an existing dedicated key comes from a Kubernetes Secret, preserve its
valueFrom.secretKeyRef entry in the combined values. A non-empty key check only
proves presence, not continuity with previously stored data.
PRELOOP_URL must be reachable from the actual agent execution namespace or
private runner, including DNS, TLS trust and egress to its resolved destination.
The native client uploads directly to
/api/v1/flows/executions/{execution_id}/artifacts; MCP access alone does not prove
this route works. An internal ingress may need an explicit existing network-policy
rule for its namespace and HTTPS port. Do not widen agent egress merely to bypass
a failed readiness check. Check upload timeouts as well as size limits.
Use a local or staging acceptance run to verify an encrypted workspace and native
artifact row, then a fresh-container native resume for the same implementation
thread. The two-turn image smoke below proves the harness's selected-session
restore separately from the encrypted HTTP transport. Old executions without an
uploaded checkpoint cannot regain their lost native context by enabling this
setting; they require an explicit cold handoff. Disable future direct uploads by
removing this overlay or setting FLOW_ARTIFACT_DIRECT_UPLOAD=false consistently,
while preserving stored artifacts and keys for recovery.
Adopt one existing publication¶
First enable feedback on the flow and direct artifact uploads on the API and
execution workers (FLOW_ARTIFACT_DIRECT_UPLOAD=true). The worker must reach
PRELOOP_URL and share the existing encryption configuration. Retain workspace
and native artifacts for the desired repair window. A native session ID alone is
not sufficient to resume a conversation.
Use the execution detail action, or preview the original successful publishing
execution through GET /api/v1/flows/executions/{execution_id}/continuation.
The preview verifies account ownership, the tracker and current PR repository,
branch, URL and head. It also exercises review, CI and repository gate reads with
at most twelve provider requests and a 25-second deadline. It writes no thread
and dispatches no execution. feedback_readable=false means feedback read permissions
or provider availability must be fixed before adoption. A permission denial
limited to repository requirement discovery permits adoption and bounded repairs,
with a warning that readiness remains unverified. Other gate blockers
remain visible in feedback_blocked_reason and the warnings.
The preview reports native_resume_expires_at whenever native resume is
offered: the effective deadline of the saved conversation is the earlier of
the workspace snapshot and native session artifact expiries (defaults are
24 hours for workspace snapshots and 168 hours for native session
artifacts). Feedback policies may run longer than stored checkpoints, so the
preview never implies a native conversation can be resumed past that
checkpoint deadline; once either artifact expires, only a fresh
published_branch_handoff is offered.
Submit the returned head using
POST /api/v1/flows/executions/{execution_id}/continuation:
{
"recovery_mode": "published_branch_handoff",
"expected_head_sha": "<head_sha from the preview>",
"acknowledge_fresh_conversation": true
}
Use native_resume when it appears in allowed_recovery_modes; fresh-conversation
acknowledgement is then unnecessary. published_branch_handoff explicitly gives
up unavailable unpublished workspace and native conversation state and starts
from the verified published PR branch. The controller binds that permission to
the selected source execution and its reserved first repair. Trigger payloads
cannot grant the exception. Subsequent turns use their own checkpoints when
available, or a controller-authorized published-branch handoff when no native
artifact was stored.
A changed head, closed PR, disabled flow, missing checkpoint capability or unreadable provider returns HTTP 409, requiring a new preview. Repeated adoption returns the existing thread without resetting its counters, deadline or terminal state when its recorded source and recovery mode match. Adopting an already subscribed PR with a different or unrecorded mode returns HTTP 409; this endpoint does not change an existing thread's recovery authority. Adoption starts the bounded subscription; it does not immediately create an agent run. The scheduler first reconciles current trusted feedback and gates.
Local validation¶
Unit fixtures cover archive identity, traversal, symlinks, credential isolation,
expiry, scheduler policy and provider outcomes. Set
FLOW_FEEDBACK_TEST_DATABASE_URL to a disposable PostgreSQL database for the
lease, crash and concurrent-worker integration tests. The suite never substitutes
the application database for this fixture. CI explicitly opts in using its
disposable PostgreSQL service; each fixture creates and drops an isolated schema.
test_flow_feedback_lifecycle.py connects fake GitHub and GitLab HTTP responses
to real CRUD reservations and scheduler turns, covering publication-time feedback,
coalesced CI/review repair, duplicate deliveries, missing events, worker restarts,
same branch/session identity, current-head readiness and manual merge. These tests
do not run a real model or publish to a provider.
NATIVE_SESSION_IMAGE_SMOKE=1 enables immutable-image Codex/OpenCode tests with a
local deterministic model HTTP fixture. The first container seeds a fact; a
second container receives only the selected native session, resumes its explicit
ID and sends the remembered fact to the fixture. No provider credentials are
needed. Set PRELOOP_DISABLE_TELEMETRY=true for all tests.