Intake
Turns an alert or a console request into a normalized incident and screens its text for prompt injection.
- intake
UNDER THE HOOD / AI INVESTIGATION
Four specialist agents read logs, code, recent changes and past incidents in parallel. One agent proposes a diagnosis, a critic challenges it, and a person approves or rejects the exact plan. Only code acts on the decision.
THE WORKFLOW
One run: intake, four specialists in parallel, gather_findings, synthesizer, the critic loop, validate_proposal and approval_gate. After the gate it runs apply_mitigation, apply_patch and verify_outcome if approved, or record_rejection; a postmortem draft follows.
Turns an alert or a console request into a normalized incident. It re-checks redaction, keeps the operator's note apart and screens the text for prompt injection.
Re-checks redaction
Keeps the operator's note apart
Screens the text for prompt injection
log_agent reads logs and code_agent reads code. change_agent lists recent changes and history_agent searches past incidents and runbooks. Each has one read-only tool and cites evidence IDs.
log_agent and code_agent: findings, file and lines
change_agent: recent commits and deploys
history_agent: past incidents and runbooks
gather_findings collects the specialists' results, keyed by specialist. The synthesizer writes the diagnosis and saves it to state as the proposal. Its default model is Gemini 3.8 Flash.
Each specialist returns success, partial or failed
A specialist past its deadline is cut off; the join still fires
synthesizer: Diagnosis output schema, no tools
Before the gate, the critic loop may revise the diagnosis and validate_proposal checks it. The gate is the single human interrupt; it writes approved or rejected to state and routes on it.
Writes approved or rejected to state
Routes on that outcome
Reruns on resume
If approved, apply_mitigation and apply_patch re-check the approval and call action-runner, which makes the writes. If rejected, record_rejection records the reviewer's comment.
Approved: each action re-checks the approval
action-runner applies the mitigation and opens the pull request
Rejected: record_rejection records the reviewer's comment
Ordered sequence · evidence, review and outcomes remain visible.
Evidence / agent workWaiting for a personFailingRecovered
The diagram draws the nodes on the main path. The step cards and the full node list below show all of them, including follow_up_agent, reviser, critic_2, the two route_critique nodes, verify_outcome and store_postmortem_draft.
The two Phase 1 live runs ended needs_more_data and did not name the connection leak. On staging since then, live runs found it and proposed the try/finally fix, and in five of six critic-loop runs the critic triggered a revision. These are a handful of runs, not a rate. Every phase that added nodes was a graph change, and unfinished runs were drained first.
EACH STEP
Every node below is in today's graph. Agents read; function nodes route and act.
Turns an alert or a console request into a normalized incident and screens its text for prompt injection.
Four agents read logs, code, recent changes and past incidents in parallel, and the join collects their results.
One agent writes the diagnosis; up to two critic passes challenge it, with at most one evidence follow-up and one revision.
Code checks the proposal first, then the run pauses until a person approves or rejects the exact plan.
Code re-checks the approval before each action, calls action-runner and checks recovery, or records why the run stopped; a postmortem draft follows.
| Node | Kind | What it does | Added in |
|---|---|---|---|
| intake | Function | Turns an alert or a console request into a normalized incident. It re-checks redaction, keeps the operator's note apart, and screens the text for prompt injection; flagged text is labelled as data. | Phase 1 |
| log_agent | Agent, one read-only tool (fetch_service_logs) | Reads logs and returns findings with evidence IDs. | Phase 1 |
| code_agent | Agent, one read-only tool | Reads code and returns the file, lines and suspected defect, with evidence IDs. | Phase 1 |
| change_agent | Agent, one read-only tool (get_recent_changes) | Lists recent commits, config diffs and deployments near the incident start. It reports timing, not blame. | Phase 2 |
| history_agent | Agent, one retrieval tool (search_knowledge) | Finds similar past incidents, runbook steps and framework guidance, cited by kb: evidence IDs, and says what matches, what differs and which remedy failed before. | Phase 4a |
| gather_findings | JoinNode | Collects the specialists' results, keyed by specialist. Each specialist returns success, partial or failed. A specialist that misses its deadline or raises an ordinary error is reported failed, so the join still fires. | Phase 1 |
| synthesizer | Agent, Diagnosis output schema, no tools | Writes the diagnosis and saves it to state as the proposal. Its default model is Gemini 3.8 Flash. | Phase 1 |
| critic_1 / critic_2 | Agent, Critique output schema, no tools | Challenge the diagnosis with typed objections and questions for missing evidence. | Phase 4b |
| route_critique_1 | Function | Saves the critique and routes to pass or revise. | Phase 4b |
| follow_up_agent | Agent, read-only tools | Answers only the critic's missing-evidence questions, with new evidence IDs and at most three tool calls. | Phase 4b |
| reviser | Agent, Diagnosis output schema, no tools | Revises the diagnosis using the new evidence and the critique. | Phase 4b |
| route_critique_2 | Function | Routes to pass or unresolved. Unresolved runs end without any action and show the open objections. | Phase 4b |
| validate_proposal | Function | Deterministic checks before the gate: the mitigation is in the allowlisted catalog with bounded parameters, the patch applies to the base commit and touches only allowlisted files. It also checks that every cited evidence ID, including kb: IDs, was issued to this run, and that flagged evidence does not support the action. | Phase 2 |
| approval_gate | Function, rerun_on_resume | The single human interrupt. It writes the outcome (approved or rejected) to state and routes on it. | Phase 1 |
| record_blocked | Function | Ends the run blocked when an action's approval re-check fails, with the reason. | Phase 2 |
| apply_mitigation | Function, with retries and a timeout | Re-checks the approval, then calls action-runner to apply an approved mitigation, if the plan has one. | Phase 2 |
| apply_patch | Function, with retries and a timeout | Re-checks the approval, then calls action-runner to open the pull request, if the plan has a change. Scripted runs use a mock and recorded replays take no action; live-tool runs open a real pull request. | Phase 1 |
| verify_outcome | Function | Polls service health for recovery within a deadline and records recovered or not recovered, with the pull request's state. It never merges. | Phase 2 |
| record_rejection | Function | Records the reviewer's comment. | Phase 1 |
| report_unresolved | Function | Records an unresolved outcome: more data needed (with the next checks and a suggested owner), open objections or a failed validation. | Phase 2 |
| postmortem_agent | Agent, no tools | Drafts the postmortem in Markdown with a low-cost model. | Phase 4a |
| route_postmortem | Function | Sends a finished run to the postmortem draft, except recorded replays and comparison members. | Phase 4a |
| store_postmortem_draft | Function | Saves the draft as a knowledge document with status draft. It is not searchable until an approver verifies it. | Phase 4a |
DESIGN RULES
Rules that keep the graph replayable and the risky parts in code.
Built
The critic loop sits before the gate and has no request for input inside it, which avoids the replay bugs seen with repeated interrupts. "Request changes" means reject with a comment; an operator can then continue it as a new linked run.
Built
The loop is drawn out in full: at most one evidence follow-up, one revision and two critic passes. Each drawn node maps one to one to the live graph in the UI, and the loop cannot run forever.
Built
An agent node returns content, and its output replaces its input on the next edge. So a route comes only from a function node, and the proposal travels past the critics through state.
Built
The same graph is used for every run, and model choice is data, not graph shape. Each run records its workflow version. If a deploy changes the graph, a run started on the old graph ends failed with a stale-workflow message and makes no model call. The version hashes only node names, kinds and edges, so before a deploy that changes the graph the operator drains unfinished runs by hand.
Built
The join fires even if a specialist fails. A specialist that misses its deadline or raises an ordinary error is reported failed and the others continue. A budget stop or a run failure still ends the run.
Built
Every action goes through write_operations with a durable key (run_id:proposal_digest:action), an exclusive claim and reconciliation after an ambiguous response. Mitigations are written as "set to value", so repeating one is harmless.
Built
Models only pick a catalog action and parameters within bounds. validate_proposal checks them, a person approves the exact plan, and only action-runner executes it.
THE APPROVAL GATE
A person approves the exact proposal they were shown, as a named reviewer the server identified. The run stays paused until then.
Built
Built
MODEL ROUTING
Every role's model is wrapped in RoutedModel, which reserves budget before each attempt. A profile gives each role a chain of candidates; if one fails, the next answers and the run shows the fallback. A comparison replays one run's frozen evidence under several profiles.
| Role | Default in the plan | Mixed profile (staging default) |
|---|---|---|
| log_agent | A specialist model, which varies by profile | OpenAI gpt-6-luna, then Gemini |
| code_agent | A specialist model, which varies by profile | Gemini 3.8 Flash |
| change_agent / history_agent | A specialist model, which varies by profile | Qwen 3 on Vertex AI, then Gemini |
| synthesizer | Gemini 3.8 Flash | Gemini 3.8 Flash |
| critic_1 / critic_2 | A reasoning model (plan section 5.2) | OpenAI gpt-6.1-sol, then Gemini |
| follow_up_agent | A model that the spike shows can make sequential tool calls | OpenAI gpt-6-luna, then Gemini |
| reviser | Gemini 3.8 Flash | Gemini 3.8 Flash |
| postmortem_agent | A low-cost model | Qwen 3 on Vertex AI, then Gemini |
WHAT YOU SEE
An incident console, a live graph with a node inspector, the run graph after the run ends, a thread with one card per agent, and an Approval screen with the hypothesis board, the critic's review and the diff. Each run's page shows its estimated model cost, and comparisons show profiles side by side.
NEXT
The durable execution page shows how a run belongs to the database, not to a browser tab, a request or an instance, and how a person's decision resumes it.
Next: Recover interrupted work →