Under the hoodagent-workflow
← Architecture showcase

UNDER THE HOOD / AI INVESTIGATION

An alert arrives.
Which agents run?

Four specialist agents read logs, code, recent changes and past incidents in parallel. One agent proposes a diagnosis, a critic challenges it, and a person approves or rejects the exact plan. Only code acts on the decision.

THE WORKFLOW

Which nodes run, and in what order?

One run: intake, four specialists in parallel, gather_findings, synthesizer, the critic loop, validate_proposal and approval_gate. After the gate it runs apply_mitigation, apply_patch and verify_outcome if approved, or record_rejection; a postmortem draft follows.

ILLUSTRATIVE SEQUENCE01 / 05
An alert arrives. Which nodes run, and in what order?
intakelog_agentcode_agentchange_agenthistory_agentgather_findingssynthesizercritic loopvalidate_proposalreport_unresolvedapproval_gatewaiting for a personapprovedapply_mitigationapply_patchif approvedrecord_rejectionif rejectedpostmortem_agentin parallel
intakeIn parallellog_agentcode_agentchange_agenthistory_agentgather_findingssynthesizerBefore the gatecritic loopvalidate_proposalapproval_gatewaiting for a personapprovedOutcomesapply_patchif approvedrecord_rejectionif rejectedapply_mitigationpostmortem_agentreport_unresolved
FUNCTION · intake

Turn the alert into a normalized incident.

Turns an alert or a console request into a normalized incident. It re-checks redaction, keeps the operator's note apart and screens the text for prompt injection.

WHAT INTAKE DOES

Re-checks redaction

Keeps the operator's note apart

Screens the text for prompt injection

AGENTS · log_agent, code_agent, change_agent, history_agent

Read the evidence in parallel.

log_agent reads logs and code_agent reads code. change_agent lists recent changes and history_agent searches past incidents and runbooks. Each has one read-only tool and cites evidence IDs.

WHAT EACH RETURNS

log_agent and code_agent: findings, file and lines

change_agent: recent commits and deploys

history_agent: past incidents and runbooks

JOIN + AGENT · gather_findings, synthesizer

Join the findings and write the diagnosis.

gather_findings collects the specialists' results, keyed by specialist. The synthesizer writes the diagnosis and saves it to state as the proposal. Its default model is Gemini 3.8 Flash.

THE JOIN AND THE PROPOSAL

Each specialist returns success, partial or failed

A specialist past its deadline is cut off; the join still fires

synthesizer: Diagnosis output schema, no tools

HUMAN INTERRUPT · approval_gate

Challenge, check, then stop for a person.

Before the gate, the critic loop may revise the diagnosis and validate_proposal checks it. The gate is the single human interrupt; it writes approved or rejected to state and routes on it.

WHAT THE GATE DOES

Writes approved or rejected to state

Routes on that outcome

Reruns on resume

FUNCTIONS · apply_mitigation, apply_patch, record_rejection

Act on the decision.

If approved, apply_mitigation and apply_patch re-check the approval and call action-runner, which makes the writes. If rejected, record_rejection records the reviewer's comment.

AFTER THE GATE

Approved: each action re-checks the approval

action-runner applies the mitigation and opens the pull request

Rejected: record_rejection records the reviewer's comment

Ordered sequence · evidence, review and outcomes remain visible.

Evidence / agent workWaiting for a personFailingRecovered

The diagram draws the nodes on the main path. The step cards and the full node list below show all of them, including follow_up_agent, reviser, critic_2, the two route_critique nodes, verify_outcome and store_postmortem_draft.

What the live runs showed

The two Phase 1 live runs ended needs_more_data and did not name the connection leak. On staging since then, live runs found it and proposed the try/finally fix, and in five of six critic-loop runs the critic triggered a revision. These are a handful of runs, not a rate. Every phase that added nodes was a graph change, and unfinished runs were drained first.

EACH STEP

What does each step do?

Every node below is in today's graph. Agents read; function nodes route and act.

Intake

Turns an alert or a console request into a normalized incident and screens its text for prompt injection.

  • intake

Specialists

Four agents read logs, code, recent changes and past incidents in parallel, and the join collects their results.

  • log_agent
  • code_agent
  • change_agent
  • history_agent
  • gather_findings

Diagnosis and critique

One agent writes the diagnosis; up to two critic passes challenge it, with at most one evidence follow-up and one revision.

  • synthesizer
  • critic_1 / critic_2
  • route_critique_1
  • follow_up_agent
  • reviser
  • route_critique_2

Approval

Code checks the proposal first, then the run pauses until a person approves or rejects the exact plan.

  • validate_proposal
  • approval_gate

Actions and record

Code re-checks the approval before each action, calls action-runner and checks recovery, or records why the run stopped; a postmortem draft follows.

  • apply_mitigation
  • apply_patch
  • verify_outcome
  • record_blocked
  • record_rejection
  • report_unresolved
  • postmortem_agent
  • route_postmortem
  • store_postmortem_draft
All nodes, with the phase that added each
Every node in the graph: its kind, its role and the phase that added it
NodeKindWhat it doesAdded in
intakeFunctionTurns an alert or a console request into a normalized incident. It re-checks redaction, keeps the operator's note apart, and screens the text for prompt injection; flagged text is labelled as data.Phase 1
log_agentAgent, one read-only tool (fetch_service_logs)Reads logs and returns findings with evidence IDs.Phase 1
code_agentAgent, one read-only toolReads code and returns the file, lines and suspected defect, with evidence IDs.Phase 1
change_agentAgent, one read-only tool (get_recent_changes)Lists recent commits, config diffs and deployments near the incident start. It reports timing, not blame.Phase 2
history_agentAgent, one retrieval tool (search_knowledge)Finds similar past incidents, runbook steps and framework guidance, cited by kb: evidence IDs, and says what matches, what differs and which remedy failed before.Phase 4a
gather_findingsJoinNodeCollects the specialists' results, keyed by specialist. Each specialist returns success, partial or failed. A specialist that misses its deadline or raises an ordinary error is reported failed, so the join still fires.Phase 1
synthesizerAgent, Diagnosis output schema, no toolsWrites the diagnosis and saves it to state as the proposal. Its default model is Gemini 3.8 Flash.Phase 1
critic_1 / critic_2Agent, Critique output schema, no toolsChallenge the diagnosis with typed objections and questions for missing evidence.Phase 4b
route_critique_1FunctionSaves the critique and routes to pass or revise.Phase 4b
follow_up_agentAgent, read-only toolsAnswers only the critic's missing-evidence questions, with new evidence IDs and at most three tool calls.Phase 4b
reviserAgent, Diagnosis output schema, no toolsRevises the diagnosis using the new evidence and the critique.Phase 4b
route_critique_2FunctionRoutes to pass or unresolved. Unresolved runs end without any action and show the open objections.Phase 4b
validate_proposalFunctionDeterministic checks before the gate: the mitigation is in the allowlisted catalog with bounded parameters, the patch applies to the base commit and touches only allowlisted files. It also checks that every cited evidence ID, including kb: IDs, was issued to this run, and that flagged evidence does not support the action.Phase 2
approval_gateFunction, rerun_on_resumeThe single human interrupt. It writes the outcome (approved or rejected) to state and routes on it.Phase 1
record_blockedFunctionEnds the run blocked when an action's approval re-check fails, with the reason.Phase 2
apply_mitigationFunction, with retries and a timeoutRe-checks the approval, then calls action-runner to apply an approved mitigation, if the plan has one.Phase 2
apply_patchFunction, with retries and a timeoutRe-checks the approval, then calls action-runner to open the pull request, if the plan has a change. Scripted runs use a mock and recorded replays take no action; live-tool runs open a real pull request.Phase 1
verify_outcomeFunctionPolls service health for recovery within a deadline and records recovered or not recovered, with the pull request's state. It never merges.Phase 2
record_rejectionFunctionRecords the reviewer's comment.Phase 1
report_unresolvedFunctionRecords an unresolved outcome: more data needed (with the next checks and a suggested owner), open objections or a failed validation.Phase 2
postmortem_agentAgent, no toolsDrafts the postmortem in Markdown with a low-cost model.Phase 4a
route_postmortemFunctionSends a finished run to the postmortem draft, except recorded replays and comparison members.Phase 4a
store_postmortem_draftFunctionSaves the draft as a knowledge document with status draft. It is not searchable until an approver verifies it.Phase 4a

DESIGN RULES

Why does routing live in code?

Rules that keep the graph replayable and the risky parts in code.

Built

One human interrupt per run

The critic loop sits before the gate and has no request for input inside it, which avoids the replay bugs seen with repeated interrupts. "Request changes" means reject with a comment; an operator can then continue it as a new linked run.

Built

Bounded critique by construction

The loop is drawn out in full: at most one evidence follow-up, one revision and two critic passes. Each drawn node maps one to one to the live graph in the UI, and the loop cannot run forever.

Built

Routing happens in function nodes

An agent node returns content, and its output replaces its input on the next edge. So a route comes only from a function node, and the proposal travels past the critics through state.

Built

Fixed topology

The same graph is used for every run, and model choice is data, not graph shape. Each run records its workflow version. If a deploy changes the graph, a run started on the old graph ends failed with a stale-workflow message and makes no model call. The version hashes only node names, kinds and edges, so before a deploy that changes the graph the operator drains unfinished runs by hand.

Built

Join survives a failure

The join fires even if a specialist fails. A specialist that misses its deadline or raises an ordinary error is reported failed and the others continue. A budget stop or a run failure still ends the run.

Built

At-least-once side effects

Every action goes through write_operations with a durable key (run_id:proposal_digest:action), an exclusive claim and reconciliation after an ambiguous response. Mitigations are written as "set to value", so repeating one is harmless.

Built

Privilege stays in code

Models only pick a catalog action and parameters within bounds. validate_proposal checks them, a person approves the exact plan, and only action-runner executes it.

THE APPROVAL GATE

What exactly does a person approve?

A person approves the exact proposal they were shown, as a named reviewer the server identified. The run stays paused until then.

Built

At the gate

  • The graph stops at approval_gate with one interrupt, and the run is paused.
  • The decision is bound to a SHA-256 digest of what the decision brief shows: the action kind, the evidence IDs, the mitigation, the patch and the base commit. Rewording a hypothesis does not change it; a different action or evidence set does.
  • The decision endpoint checks the approver role. It accepts the decision only while the run is paused and the digest matches. Otherwise it returns 409 and records nothing. A second decision also gets 409.
  • The reviewer is read on the server from the verified sign-in token. The client never supplies it.
  • The decision comment is redacted before it is stored or sent to ADK.
  • The Approve and Reject buttons stay off until the run is paused.

Built

Before each live action

  • The approval is tied to the reviewer, the run and the proposal digest, and it expires after 15 minutes.
  • apply_mitigation and apply_patch re-check it just before calling action-runner: the digest, the run, the expiry and, with live tools, the deployed commit. action-runner checks the stored approval again. A failed check blocks that action and ends the run; an action that already ran stays applied.
  • One approval covers the whole action plan: the mitigation and the pull request.
  • On staging, an expired approval ended the run blocked with no write, and a changed proposal sent straight to action-runner was refused.

MODEL ROUTING

Which model answers each role, and what if it fails?

Every role's model is wrapped in RoutedModel, which reserves budget before each attempt. A profile gives each role a chain of candidates; if one fails, the next answers and the run shows the fallback. A comparison replays one run's frozen evidence under several profiles.

The default model for each role in the plan, and its chain in the mixed profile
RoleDefault in the planMixed profile (staging default)
log_agentA specialist model, which varies by profileOpenAI gpt-6-luna, then Gemini
code_agentA specialist model, which varies by profileGemini 3.8 Flash
change_agent / history_agentA specialist model, which varies by profileQwen 3 on Vertex AI, then Gemini
synthesizerGemini 3.8 FlashGemini 3.8 Flash
critic_1 / critic_2A reasoning model (plan section 5.2)OpenAI gpt-6.1-sol, then Gemini
follow_up_agentA model that the spike shows can make sequential tool callsOpenAI gpt-6-luna, then Gemini
reviserGemini 3.8 FlashGemini 3.8 Flash
postmortem_agentA low-cost modelQwen 3 on Vertex AI, then Gemini

WHAT YOU SEE

What does the engineer see?

An incident console, a live graph with a node inspector, the run graph after the run ends, a thread with one card per agent, and an Approval screen with the hypothesis board, the critic's review and the diff. Each run's page shows its estimated model cost, and comparisons show profiles side by side.

NEXT

Follow a run through an interruption.

The durable execution page shows how a run belongs to the database, not to a browser tab, a request or an instance, and how a person's decision resumes it.

Next: Recover interrupted work →