Each investigation is one row in the database with one status. A queue runs it, a lease says who is working on it, and a sweeper picks up work that lost its worker. A paused run waits for a person, even across a restart or a new revision.
Every run is one row with one status. These are the transitions the plan allows. Each write is a compare-and-set on that row, and a status is written only by the party the table names.
ILLUSTRATIVE SEQUENCE01 / 07
One run, one status. Which transitions does the plan allow?
EXECUTE HANDLER
queued → running
The execute handler takes the lease. Only one executor holds a run's lease at a time.
WHO WRITES IT
The execute handler, while it holds the lease
Each write is a compare-and-set on the run row
EXECUTE HANDLER · APPROVAL GATE
running → paused
The graph stops at the approval gate. Only the executor holding the current lease token can write this.
WHO WRITES IT
The execute handler, when the graph stops at the approval gate
Only the current lease token can write this
DECISION ENDPOINT
paused → queued
The decision endpoint records a decision on the exact proposal digest, in one compare-and-set. It accepts it only while the run is paused and the digest matches; otherwise it returns 409 and records nothing.
ONE COMPARE-AND-SET
Accepted only while the run is paused
and only if the proposal digest matches
Otherwise 409, and nothing is recorded
EXECUTOR
running → completed / rejected / unresolved
The workflow returns with outcome approved, rejected or unresolved. The executor maps it to a status only when the whole workflow has returned.
OUTCOME TO STATUS
approved → completed
rejected → rejected
unresolved → unresolved
EXECUTOR · ROUTEDMODEL
running → budget_stopped
RoutedModel cannot reserve budget for a model call. Recorded outside the graph.
TERMINAL
Recorded outside the graph
A terminal status releases the lease
EXECUTOR
running → failed
A non-retryable error, or the attempt cap is passed.
TERMINAL
A non-retryable error
Or the attempt cap is passed
A terminal status releases the lease
ADMIN ENVIRONMENT RESET
any unfinished status → cancelled
The admin environment reset cancels any unfinished run.
TERMINAL
Any unfinished status
Set by the admin environment reset
Ordered sequence · evidence, review and outcomes remain visible.
Evidence / agent workWaiting for a personFailingRecovered
Lease expiry is not a transition. When a lease expires the sweeper only enqueues a task, and a new executor takes a new lease.
Run statuses, whether each is terminal, and what sets it
Status
Terminal
Set by
queued
no
POST /incidents; the decision endpoint (from paused only)
running
no
The execute handler, while it holds the lease
paused
no
The execute handler, when the graph stops at the approval gate
completed
yes
The executor, when the workflow returns with outcome = approved
rejected
yes
The executor, when the workflow returns with outcome = rejected
unresolved
yes
The executor, when the workflow returns with outcome = unresolved
budget_stopped
yes
The executor, when RoutedModel cannot reserve budget for a model call
failed
yes
The executor, on a non-retryable error or past the attempt cap
cancelled
yes
The admin environment reset, for any unfinished run
A terminal status releases the lease and ends the sweeper's interest in the run. A budget stop after approval is terminal too: actions already taken stay recorded in write_operations, and the postmortem is not drafted.
LEASE AND SWEEPER
Who is working on a run right now?
A worker crash, in order. Only one executor holds a run's lease at a time, and the lease token is what the database checks on every executor status write and event projection to runs and run_events.
01 / Cloud Tasks → execute endpoint
Cloud Tasks calls the private execute endpoint. It accepts only Cloud Tasks' OIDC token for the executor service account (issuer, audience and account checked). Browser Firebase tokens are rejected there.
02 / executor A → runs row
The handler takes the execution lease on the run (one active executor per run) with a fresh lease token, and increments attempts.
It returns 200 at once, without starting the runner, when the run is paused, terminal, or leased by a live executor. Past the attempt cap it sets failed instead of running.
03 / executor A → runs, run_events
It renews the lease while the run is running. Every status write and event projection is conditional on its token, and a failed renewal stops the executor at once.
04 / executor A ×
The executor dies. Its lease stops being renewed and then expires.
05 / sweeper → Cloud Tasks
The sweeper, called by Cloud Scheduler, finds runs that are running with an expired lease, or queued with no live lease at least a minute after they were queued or decided. It re-enqueues them and stops at the attempt cap.
It only enqueues; it never runs the graph. It never touches a paused or terminal run.
06 / executor B → runs row
A new executor takes a new lease token and resumes the interrupted invocation by invocation_id. Only nodes with no recorded output rerun.
07 / executor A (stale) → runs, run_events
If the stale executor wakes up, its status and event writes fail on its old token, so it cannot change runs or run_events after another executor has taken over.
Limit:ADK session appends are not fenced by the lease token. ADK's own revision check (StaleSessionError in DatabaseSessionService.append_event) refuses the stale writer's held append, and its attempt ends with a lost lease. A local takeover test shows the run still reaching its normal outcome.
On staging, live updates reach the browser as a stream, not in delayed batches, and a reconnect resumes from the last event it received.
What happens when a worker dies or a person comes back later?
They are kept separate. Both end with a task on the queue and an executor that skips what is already done, but they start from different events and ADK is told different things.
Human decision
Trigger: a person approves or rejects the exact proposal.
triage-api
Checks the approver role, binds the decision to the proposal digest, and records it in the same compare-and-set that moves the run paused → queued. The server adds the reviewer identity; the client never supplies it.
Cloud Tasks
A task for the private execute endpoint is enqueued, on whichever revision is serving.
executor
Takes the lease. A decision counts as delivered when the adk_request_input function response for the gate's interrupt is in the run's ADK session, read from the session, not from a flag. If it is recorded but not delivered, the executor sends it as the function response.
ADK Runner
Matches the response id to the paused invocation, so no invocation_id is passed. Completed nodes replay from events without new model calls. If the response is already delivered, the executor resumes by invocation_id and never sends it again, because an unmatched response starts a new run.
Executor crash mid-run
Trigger: the executor dies and its lease expires.
Cloud Tasks or sweeper
Cloud Tasks retries the task, or the sweeper re-enqueues it.
new executor
Takes a new lease token and resumes the interrupted invocation by invocation_id, as ADK's resume docs describe.
ADK Runner
Completed nodes of the static graph are skipped. Only nodes with no recorded output rerun (ADR 0001, S2).
None of this is exactly-once execution. Nodes and model calls can repeat; only external writes are idempotent, each keyed by run, proposal and action in write_operations. The decision records hold the evidence; each card says what it shows and what it does not.
Executor dies mid-run
Resumes by invocation_id; only nodes with no recorded output rerun.
Evidence: Verified: ADR 0001 S2 rows (a model error, a cancelled task, and a killed process on Omni).
Limit: Nodes and model calls can repeat; only external writes are idempotent, through write_operations.
Provenresume without rerunning finished nodes
One specialist fails
The join still fires; that specialist is reported failed and the others continue.
Evidence: Verified for a missed deadline (guard node, ADR 0001) and for an error raised inside a specialist (guard test, ADR 0009); on staging a specialist's model timeout was reported failed and the run went on (ADR 0010).
Provenmissed deadline
Provenerror inside a specialist
Approval after a deploy
Compatible deploy: the run resumes on the new revision. Changed graph: the run ends failed without a model call.
Evidence: Verified: a paused run resumed and completed on a new Cloud Run revision (ADR 0001); stale-version failure (tests/runs/test_executor.py).
Limit: A restart onto a new revision is proven; a deploy with changed code is untested end to end.
Provenrestart onto a new revision
Provenversion mismatch fails closed
Takeover during a save
Takeover while an ADK event is being saved. The run reaches the outcome it would have reached without the takeover; its session stays consistent and replays without a repeated decision; external effects are reconciled. It recovers within the attempt cap.
Evidence: Verified locally (ADR 0011): with a takeover while a specialist event or apply_patch was being saved, the run paused once or completed with one pull request, and the stale writer's append was refused.
Limit: Local only, on Postgres with fake GitHub and storefront services; no takeover was forced on staging.
Proventakeover test
Response lost after a write
An external write succeeds but its response is lost. Reconciled, not repeated.
Evidence: Verified on staging (ADR 0011): each action's first answer was dropped on purpose, and the retry found the existing pull request and pool-limit change instead of repeating them. Locally, a lost answer at three points each ended with one branch, one pull request and one mitigation.
Limit: The lost answer is a staging test setting, not a real network drop.