Under the hooddurable-execution
← Architecture showcase

UNDER THE HOOD / RECOVER INTERRUPTED WORK

Close the tab.
The investigation keeps going.

Each investigation is one row in the database with one status. A queue runs it, a lease says who is working on it, and a sweeper picks up work that lost its worker. A paused run waits for a person, even across a restart or a new revision.

RUN STATES

One run, one status.

Every run is one row with one status. These are the transitions the plan allows. Each write is a compare-and-set on that row, and a status is written only by the party the table names.

ILLUSTRATIVE SEQUENCE01 / 07
One run, one status. Which transitions does the plan allow?
any unfinished statusqueuedrunningpausedcompletedrejectedunresolvedbudget_stoppedfailedcancelled1234567
EXECUTE HANDLER

queued → running

The execute handler takes the lease. Only one executor holds a run's lease at a time.

WHO WRITES IT

The execute handler, while it holds the lease

Each write is a compare-and-set on the run row

EXECUTE HANDLER · APPROVAL GATE

running → paused

The graph stops at the approval gate. Only the executor holding the current lease token can write this.

WHO WRITES IT

The execute handler, when the graph stops at the approval gate

Only the current lease token can write this

DECISION ENDPOINT

paused → queued

The decision endpoint records a decision on the exact proposal digest, in one compare-and-set. It accepts it only while the run is paused and the digest matches; otherwise it returns 409 and records nothing.

ONE COMPARE-AND-SET

Accepted only while the run is paused

and only if the proposal digest matches

Otherwise 409, and nothing is recorded

EXECUTOR

running → completed / rejected / unresolved

The workflow returns with outcome approved, rejected or unresolved. The executor maps it to a status only when the whole workflow has returned.

OUTCOME TO STATUS

approved → completed

rejected → rejected

unresolved → unresolved

EXECUTOR · ROUTEDMODEL

running → budget_stopped

RoutedModel cannot reserve budget for a model call. Recorded outside the graph.

TERMINAL

Recorded outside the graph

A terminal status releases the lease

EXECUTOR

running → failed

A non-retryable error, or the attempt cap is passed.

TERMINAL

A non-retryable error

Or the attempt cap is passed

A terminal status releases the lease

ADMIN ENVIRONMENT RESET

any unfinished status → cancelled

The admin environment reset cancels any unfinished run.

TERMINAL

Any unfinished status

Set by the admin environment reset

Ordered sequence · evidence, review and outcomes remain visible.

Evidence / agent workWaiting for a personFailingRecovered

Lease expiry is not a transition. When a lease expires the sweeper only enqueues a task, and a new executor takes a new lease.

Run statuses, whether each is terminal, and what sets it
StatusTerminalSet by
queuednoPOST /incidents; the decision endpoint (from paused only)
runningnoThe execute handler, while it holds the lease
pausednoThe execute handler, when the graph stops at the approval gate
completedyesThe executor, when the workflow returns with outcome = approved
rejectedyesThe executor, when the workflow returns with outcome = rejected
unresolvedyesThe executor, when the workflow returns with outcome = unresolved
budget_stoppedyesThe executor, when RoutedModel cannot reserve budget for a model call
failedyesThe executor, on a non-retryable error or past the attempt cap
cancelledyesThe admin environment reset, for any unfinished run

A terminal status releases the lease and ends the sweeper's interest in the run. A budget stop after approval is terminal too: actions already taken stay recorded in write_operations, and the postmortem is not drafted.

LEASE AND SWEEPER

Who is working on a run right now?

A worker crash, in order. Only one executor holds a run's lease at a time, and the lease token is what the database checks on every executor status write and event projection to runs and run_events.

  1. 01 / Cloud Tasks → execute endpoint

    Cloud Tasks calls the private execute endpoint. It accepts only Cloud Tasks' OIDC token for the executor service account (issuer, audience and account checked). Browser Firebase tokens are rejected there.

  2. 02 / executor A → runs row

    The handler takes the execution lease on the run (one active executor per run) with a fresh lease token, and increments attempts.

    It returns 200 at once, without starting the runner, when the run is paused, terminal, or leased by a live executor. Past the attempt cap it sets failed instead of running.

  3. 03 / executor A → runs, run_events

    It renews the lease while the run is running. Every status write and event projection is conditional on its token, and a failed renewal stops the executor at once.

  4. 04 / executor A ×

    The executor dies. Its lease stops being renewed and then expires.

  5. 05 / sweeper → Cloud Tasks

    The sweeper, called by Cloud Scheduler, finds runs that are running with an expired lease, or queued with no live lease at least a minute after they were queued or decided. It re-enqueues them and stops at the attempt cap.

    It only enqueues; it never runs the graph. It never touches a paused or terminal run.

  6. 06 / executor B → runs row

    A new executor takes a new lease token and resumes the interrupted invocation by invocation_id. Only nodes with no recorded output rerun.

  7. 07 / executor A (stale) → runs, run_events

    If the stale executor wakes up, its status and event writes fail on its old token, so it cannot change runs or run_events after another executor has taken over.

    Limit: ADK session appends are not fenced by the lease token. ADK's own revision check (StaleSessionError in DatabaseSessionService.append_event) refuses the stale writer's held append, and its attempt ends with a lost lease. A local takeover test shows the run still reaching its normal outcome.

On staging, live updates reach the browser as a stream, not in delayed batches, and a reconnect resumes from the last event it received.

TWO RESUME PATHS

What happens when a worker dies or a person comes back later?

They are kept separate. Both end with a task on the queue and an executor that skips what is already done, but they start from different events and ADK is told different things.

Human decision

Trigger: a person approves or rejects the exact proposal.

  1. triage-api

    Checks the approver role, binds the decision to the proposal digest, and records it in the same compare-and-set that moves the run paused → queued. The server adds the reviewer identity; the client never supplies it.

  2. Cloud Tasks

    A task for the private execute endpoint is enqueued, on whichever revision is serving.

  3. executor

    Takes the lease. A decision counts as delivered when the adk_request_input function response for the gate's interrupt is in the run's ADK session, read from the session, not from a flag. If it is recorded but not delivered, the executor sends it as the function response.

  4. ADK Runner

    Matches the response id to the paused invocation, so no invocation_id is passed. Completed nodes replay from events without new model calls. If the response is already delivered, the executor resumes by invocation_id and never sends it again, because an unmatched response starts a new run.

Executor crash mid-run

Trigger: the executor dies and its lease expires.

  1. Cloud Tasks or sweeper

    Cloud Tasks retries the task, or the sweeper re-enqueues it.

  2. new executor

    Takes a new lease token and resumes the interrupted invocation by invocation_id, as ADK's resume docs describe.

  3. ADK Runner

    Completed nodes of the static graph are skipped. Only nodes with no recorded output rerun (ADR 0001, S2).

FAILURE SCENARIOS

What can go wrong, and what happens next?

None of this is exactly-once execution. Nodes and model calls can repeat; only external writes are idempotent, each keyed by run, proposal and action in write_operations. The decision records hold the evidence; each card says what it shows and what it does not.

Executor dies mid-run

Resumes by invocation_id; only nodes with no recorded output rerun.

Evidence: Verified: ADR 0001 S2 rows (a model error, a cancelled task, and a killed process on Omni).

Limit: Nodes and model calls can repeat; only external writes are idempotent, through write_operations.

  • Provenresume without rerunning finished nodes

One specialist fails

The join still fires; that specialist is reported failed and the others continue.

Evidence: Verified for a missed deadline (guard node, ADR 0001) and for an error raised inside a specialist (guard test, ADR 0009); on staging a specialist's model timeout was reported failed and the run went on (ADR 0010).

  • Provenmissed deadline
  • Provenerror inside a specialist

Approval after a deploy

Compatible deploy: the run resumes on the new revision. Changed graph: the run ends failed without a model call.

Evidence: Verified: a paused run resumed and completed on a new Cloud Run revision (ADR 0001); stale-version failure (tests/runs/test_executor.py).

Limit: A restart onto a new revision is proven; a deploy with changed code is untested end to end.

  • Provenrestart onto a new revision
  • Provenversion mismatch fails closed

Takeover during a save

Takeover while an ADK event is being saved. The run reaches the outcome it would have reached without the takeover; its session stays consistent and replays without a repeated decision; external effects are reconciled. It recovers within the attempt cap.

Evidence: Verified locally (ADR 0011): with a takeover while a specialist event or apply_patch was being saved, the run paused once or completed with one pull request, and the stale writer's append was refused.

Limit: Local only, on Postgres with fake GitHub and storefront services; no takeover was forced on staging.

  • Proventakeover test

Response lost after a write

An external write succeeds but its response is lost. Reconciled, not repeated.

Evidence: Verified on staging (ADR 0011): each action's first answer was dropped on purpose, and the retry found the existing pull request and pool-limit change instead of repeating them. Locally, a lost answer at three points each ended with one branch, one pull request and one mitigation.

Limit: The lost answer is a staging test setting, not a real network drop.

  • Provenlost-response write test

NEXT

Follow a decision into the approval checks.

The security page shows who may approve, how an approval is tied to one proposal, and which separate service holds the write credentials.

Next: Security and approvals →