Proven
A named test, or a controller-confirmed result in a decision record, backs the claim as worded. The claim reaches no further than its wording, and its limits say where it stops.
RECORDED EVIDENCE
Every claim this site makes, with what backs it and where it stops. Proven means a named test or a controller-confirmed result in a decision record shows it. In build means it is designed and not shown yet.
HOW TO READ THIS
Each claim has a status and the evidence behind it. Open a claim in the ledger to see the test or decision record, where it was observed and where it stops.
A named test, or a controller-confirmed result in a decision record, backs the claim as worded. The claim reaches no further than its wording, and its limits say where it stops.
Designed, not shown yet. Each In build claim names the phase that will build it.
Where the evidence was observed: local, local AlloyDB Omni, CI or staging on Google Cloud. A decision record also gives the date when it has one.
THE LEDGER
Open a claim to see its evidence and its limits. Claims are grouped by the page that explains them.
With new App, Runner and session-service objects, a run paused at the approval gate resumed with the decision; afterwards only postmortem_agent called a model. Tested on SQLite and local Omni.
“new App, Runner and session service objects; after the decision only `postmortem_agent` calls a model”
Limits. Restart is simulated by building new objects, not by killing a process. The record gives no date. The cloud result is a separate claim.
A run paused at the gate stayed paused while Cloud Run revision 00002 was replaced by 00003, which served the decision; the run completed with one decision and one finish.
“a new revision (`00002` → `00003`) served the decision; the run completed with one decision and one finish”
“"revisions": [ "triage-api-00002-bsg", "triage-api-00003-hf6" ]”
“"decisions": 1, "finished": 1”
Limits. Scripted models (the longest execute was 43 s with scripted models). The record does not say what changed between revisions, so this proves a restart onto a new revision, not a deploy with changed code. The retired check script is not cited because it was deleted; its results survive in the decision record and the cloud-checks document.
After env-down then env-up, the run was still paused with its proposal digest unchanged, then completed on the decision.
“after `env-down` and `env-up` the run is still `paused` with its digest unchanged, then completes on the decision”
“"status_after_restart": "paused", "digest_unchanged": true,”
Limits. One run. The check script is retired, so this cannot be re-run from the repo.
After a model error, a cancelled task and a killed process on Omni, resuming by invocation id finished the run; only nodes with no recorded output reran.
“PASS for a model error, a cancelled task and a killed process on Omni; only nodes with no recorded output rerun”
Limits. Other deaths in the tests are injected faults (raised errors, SQL edits, an expired lease), not killed processes; the killed-process case ran on local Omni. Not cloud-tested. Not exactly-once: nodes and model calls can repeat. Every resume re-appends about 11 replayed events.
A repeated, conflicting decision leaves outcome unchanged and makes no model calls; ADK replays the gate's recorded route. The gate also ignores decisions once outcome is set.
“a repeated, conflicting decision leaves `outcome` unchanged and makes no model calls”
Limits. Not re-proven in the cloud: proven locally; the API and database logic is the same on Cloud Run.
Paused, finished or live-leased runs return skipped with no model calls. Two concurrent executes on one run gave one paused and one skipped; the synthesizer was called once.
“two concurrent executes on one run give one `paused` and one `skipped`, and the synthesizer is called once”
Limits. Local only. Cloud Tasks duplicates were not induced. Not exactly-once: after a failed attempt the synthesizer ran twice.
After a failed attempt the run is running with no lease and attempts 1; the next attempt pauses it. The endpoint returns 503, then 200 paused on the retry.
“after a failed attempt the run is `running` with no lease and `attempts` 1; the next attempt pauses it”
Limits. The failure is a raised error in a scripted model, not a killed process.
With a superseded lease token, renew, append_events, pause, finish and release all return False and write nothing to runs or run_events. ADK session appends are not fenced.
“an old token's renew, `append_events`, pause, finish and release all return False and write nothing”
Limits. Covers runs and run_events only. ADK session appends are not fenced by the lease token; ADK's own revision check refuses a stale append, which the takeover claim tests.
With a cap of 2, two failed attempts, then the third claim marks the run failed (attempt cap reached) and returns skipped with no model calls.
“the third claim marks the run `failed` (`attempts` 2, `last_error` "attempt cap reached")”
Limits. Resumes count against the cap (read from the code, not from a test assertion). A run that pauses on its last permitted attempt accepts a decision but then ends failed, undelivered: a known edge, not tested.
Runs queued past the grace period or running with an expired lease are returned; fresh, live-leased and paused runs are not. A dropped rejection is swept and ends rejected.
“`queued` past the grace period and `running` with an expired lease are returned”
Limits. Local only. The cloud side only shows that Cloud Scheduler reached the endpoint (next claim), not that a lost task was recovered on staging.
Death before delivery, right after ADK stored the decision, and in postmortem_agent each ended completed with exactly one decision event in the session.
“each case ends `completed` with one decision event in the session”
Limits. The deaths are simulated: an expired lease with no send, an error raised while projecting the decision event, and an error raised in a scripted model; none kills a process.
Spike graph only: approval gave completed, rejection rejected, needs_more_data unresolved without pausing, and a crash in postmortem_agent was recovered to completed.
“a crash in `postmortem_agent` is recovered to `completed`”
Limits. Phase 0 spike graph (tests/spike_graph.py), not the production graph. Production now also ends needs_more_data unresolved without the gate; that is tested separately.
With executor A stalled mid-append, B took the lease and finished: the run paused once, or completed with one PR. ADK refused A's stale append, ending its attempt LOST_LEASE.
“the run pauses once at the gate, then completes after approval”
“its held append is refused by ADK (`StaleSessionError`) and its attempt ends `LOST_LEASE`”
Limits. Local only, on Postgres with fake GitHub and storefront services. The stall is made inside the session service; no takeover was forced on staging. ADK's revision check did the fencing, so no change to it was needed.
On staging each action's first answer was dropped and the retry reconciled: one PR, one pool-limit change. Locally, three lost-answer points each ended with one branch, PR and mitigation.
“Each action's first answer was dropped (`503`, "ambiguous write … staging test setting") and the retry reconciled (`200`).”
“each ends with one branch, one PR, one mitigation, and two `succeeded` rows”
Limits. The lost answer is a staging test setting, not a real network drop. The staging write records are inferred from the answers, because the database is private. A definite refusal is not retried; the run ends failed.
A run whose workflow version differs from the deployed one ends failed with a stale-workflow message and makes no model call.
Limits. Backed by a CI test that passed in CI's backend job. There is no separate recorded result row. The workflow version hashes only node names, kinds and edges.
A paused run resumed and completed on a new Cloud Run revision (00002 to 00003); a later restart replaced triage-api revision 00007 with 00008 at 100% traffic.
“a new revision (`00002` → `00003`) served the decision; the run completed with one decision and one finish”
“triage-api revision `00007` → `00008`, 100% traffic”
Limits. The ADRs do not show that the revisions differed in code. The rule that prompt, UI and API changes are compatible is a design rule, not a tested deploy.
With the local server closing streams every 3 s, the page reconnected 3 times; the run's 15 cards each appeared once, with one decision and one finish.
“the run's 15 cards (from 11 at the pause) each appear once, with one decision and one finish”
Limits. Local Chromium. Deduplicating a replayed event is proven by mergeEvent's unit tests, not in the browser. The stream polls every 0.5 s, so a status shorter than one poll can be skipped. The browser tests are TypeScript and are not listed as test evidence.
A wrong digest and a second decision each get 409; exactly one approvals row exists after the second. A decision before the pause commits gets 409 and records nothing.
“A wrong digest and a second decision each get 409, and exactly one `approvals` row exists after the second”
Limits. Local only. The 409 body is the same for not paused and digest differs. The re-check before each action is a separate claim.
With a controlled stream, Approve and Reject stay disabled after the gate event while running, enable on paused, and disable on end. A real local decision completed the run.
“Approve and Reject stay disabled after the gate event and digest arrive while the run is `running`, enable on `paused`”
Limits. Local stack (jsdom and Chromium). Behaviour on staging was not separately checked. The decision record row is informational. The browser tests are TypeScript and are not listed as test evidence.
In a real Google sign-in on localhost, the approval's reviewer and the decision_recorded audit actor were the owner's verified email; the client does not supply it.
“`decision_recorded` (the owner's email, approved)”
Limits. Local only. Cloud Run sign-in was not verified at the time. The later staging sign-in was owner-observed plus an inference, so it is not claimed here.
Approvals expire after 15 minutes. Before each action, code re-checks the digest, run, expiry and deployed commit, and action-runner checks the stored approval again.
“Ended `blocked` ("the approval expired") at the re-check before `apply_mitigation`; no new branch or PR”
“`403 {"refused": "the proposal differs from the approved one"}`; still two PRs.”
Limits. On staging, one expired approval and one changed proposal were blocked. The other reasons to block are covered by tests. The deployed-commit check runs only with live tools.
approval_gate approves the combined plan. apply_mitigation and apply_patch act only after it and are each re-checked first; verify_outcome then reads the result.
“`code_patch` (try/finally in `get_cart_totals`) + `set_pool_limit` 30, 6/6 checks.”
“`apply_mitigation` and `apply_patch` each call `recheck_approval`”
Limits. The plan is approved or rejected as a whole; an approver cannot approve the mitigation and refuse the pull request.
action-runner alone reads the GitHub App key and its own database credential. Cloud Run IAM limits who can call it, and it checks the caller per route.
“service runs as the `action-runner` account, which alone reads the GitHub App key and the”
Limits. Secret access is set in Terraform; no staging test probes it. The API still writes its own database, including approval rows, so a compromised API could forge one; it holds no write credential for GitHub or the shop.
Without sign-in, triage-web answered 200 for the public page, /api/triage/me answered 401, and calling triage-api directly answered 403 from Cloud Run IAM.
“triage-web 200; `/api/triage/me` 401; triage-api direct 403 (IAM)”
Limits. A service-perimeter check. It says nothing about which database role the API uses.
In proof run 1 on staging, four refusal probes each got 403: three from the storefront's own caller checks and one from Cloud Run.
“Four refusals: 403 each (three from the storefront's caller checks, one from Cloud Run)”
Limits. The record has no date and does not name which callers were tried. It predates action-runner, so it does not show that mitigations accept only action-runner's identity.
Model tools are four reads. Models propose catalog actions and find-and-replace edits; validate_proposal checks them, a human approves, and action-runner writes.
“GitHub: one branch `triage/f706f922-f6355133`, one PR (#1, by the App).”
Limits. The app never merges a pull request. Operator-only admin routes (reset, demo commit) also write, but no model can reach them.
25 concurrent reservations against a cap that fits 10: exactly 10 succeeded. This tests reservation arithmetic, not real provider spend.
“25 concurrent reservations against a cap for 10 → exactly 10 succeed”
Limits. Reservation arithmetic only; scripted and test amounts, not real provider spend.
One plugin-visible model call per agent turn but one ledger reservation per attempt; an attempt that cannot reserve is never sent; a timed-out attempt keeps its reservation.
“an attempt that can't reserve is never sent; a timed-out attempt keeps its reservation”
Limits. Scripted attempts. RoutedModel imports ADK's private _RequestSnapshot, so an ADK upgrade must re-run these tests. Fallback chains are a separate claim.
A budget of 1 micro-USD gives status budget_stopped with no model call.
“a budget of 1 micro gives `budget_stopped` with no model call”
Limits. One scripted case. On staging a live run with a small budget also stopped, at the synthesizer.
model_calls stores each attempt's cost from a versioned price table, including fallbacks and repairs; on staging each run's summed cost matched its budget ledger.
“Every run's summed `cost_micros` equals its ledger's spent amount.”
Limits. An estimate from token counts and a price table, not the provider's invoice. A timed-out attempt keeps its reservation as uncertain cost. One run per profile was measured.
A recorded replay and a real-sign-in run each wrote run_started, gate_raised, decision_recorded and run_finished in the audit log, with the approver's email on the decision.
“Audit: `run_started`, `gate_raised`, `decision_recorded`, `run_finished`.”
“`run_started` (the owner's email), `gate_raised` (`system:executor`)”
Limits. Local runs only. Blocked actions are audited; successful writes are recorded by action-runner in its own table, not as audit rows. The log is append-only for the application, not tamper-proof.
Granting a second Google account the viewer role wrote a user_roles row and a roles_set audit row naming the owner's email.
“`user_roles` holds it, and the audit has `roles_set` by the owner's email”
Limits. Local. Role removal has no result row of its own, so it is not claimed.
A database trigger refuses UPDATE, DELETE and TRUNCATE on the audit table. The API role holds DELETE on its tables, so the trigger, not a grant, enforces this.
Limits. Backed by a CI test that passed in CI's backend job. The trigger is a design decision with no result row of its own; the CI test is the evidence. An administrator can still disable the trigger. The log is append-only for the application, not tamper-proof, and not immutable.
Planned for Phase 5: a filterable audit timeline per run with an in-app span view. Blocked actions are audit rows today; successful write results are kept in write_operations.
Limits. Not built. Successful writes are recorded by action-runner in its own table, not as audit rows.
A planted API key, email, connection-string password and AWS key appeared in no ADK session event or state, model request, run_events, runs or approvals row.
“an API key, an email, a password inside a connection string (note) and an AWS key (comment) appear in none of the ADK session events”
Limits. Not scanned: budget_ledger and ADK's other tables (app_states, user_states). Only those four secret types. Local. Tool-output redaction is a separate claim (tool-output-redaction-and-content-logging).
A provider error carrying password=... was stored in last_error and shown as [REDACTED:secret_assignment]. Only this one pattern was tested.
“A provider error carrying `password=...` is stored and shown as `[REDACTED:secret_assignment]`”
Limits. One pattern tested.
In the Phase 0 check, the incident text, a note and a decision comment each landed in BigQuery redacted: marker present, canary absent. Content logging has since been switched off.
“the incident text, a note and a decision comment each landed redacted (marker present, canary absent)”
“"incident": "redacted", "note": "redacted", "comment": "redacted"”
Limits. A Phase 0 check. Content logging was off in Phase 1 and turned on again once tool output was redacted (see the tool-output redaction claim).
On a live staging run, 66 agent_events rows were written (5 LLM_REQUEST, 5 LLM_RESPONSE, 2 TOOL_COMPLETED, others) and content and content_parts are empty on every row.
“`content` and `content_parts` are empty on every row”
Limits. One live run. Content logging was turned on in Phase 2, once tool output was redacted.
Every fetched log line, code file and commit goes through redaction before it is trimmed. With that in place, staging logs prompt and response content to BigQuery.
“81 content rows for the live run, with no secret shapes.”
Limits. Patterns catch known secret shapes only. The first staging check found one approver email in BigQuery, from the approval gate rather than a tool; it was deleted and the cause fixed.
A regex screen flags tool results and incident text. Flagged evidence is labelled as data, audited, and fails validate_proposal when cited as support.
“WARNING line was fetched (one of 72 evidence items) and flagged `ignore instructions`. The gate”
“owner `payments-team`) and did not cite the line.”
Limits. A regex screen: it can miss injections and flag harmless lines. It stops a proposal that rests on flagged evidence; it can't prove the model wasn't swayed.
Real Google sign-in on localhost via Firebase: the owner (bootstrap admin) started and approved a run to completed; GET /me without a token returned 401.
“Signed in as the project owner, a bootstrap admin. Started a scripted run, approved it, and the run reached `completed`.”
“`GET /me` → 401”
Limits. A real token refresh near expiry and the cookie round trip after an hour are unverified; the no-role person was only unit-tested. Only google.com verified-email sign-ins are accepted.
200 for the public web page, 401 at the app, 403 from IAM on direct API calls.
“triage-web 200; `/api/triage/me` 401; triage-api direct 403 (IAM)”
Limits. Says nothing about role checks on staging.
Giving a second Google account viewer wrote triage.user_roles and a roles_set audit row; bootstrap admins come from TRIAGE_ADMIN_EMAILS.
“`user_roles` holds it, and the audit has `roles_set` by the owner's email”
Limits. The viewer's actual restrictions on a second account were not exercised for real. Bootstrap admins come from a configuration setting.
A guard node returning failed handles a per-specialist deadline; ADK's own node timeout fails the whole workflow, so it is not used.
“PASS with a guard node returning `failed`; ADK's node `timeout` fails the whole workflow, so it is not used”
Limits. Covers a missed deadline only. An error raised inside a specialist is a separate claim (specialist-error-join-still-fires).
A guard node turns an ordinary error into a redacted failed result, and the join still fires. Budget stops and run failures still end the run.
“the guard reported the specialist failed and the run went on”
Limits. The one error seen on staging was a model timeout. The failed result carries a redacted reason and no findings. The guard depends on a private ADK error class.
A recorded replay paused at the gate with the same proposal digest as the live run, spend 0, and completed 10.4 s after creation. Five recordings were stored.
“Paused at the gate with the **same** proposal digest (`a24ae76c…`). Spend 0.”
Limits. Replay by a different process than the one that recorded it is not verified; it replayed in the same process. The staging replay was owner-observed and is not claimed. Recorded runs disable actions.
make api-local-live, native gemini-3.8-flash, mock tools: paused at the gate 36.4 s after intake, 5 model calls (2 per specialist plus 1 synthesis), approved, completed.
“Paused at the gate 36.4 s after intake”
Limits. One run. It ended needs_more_data with no patch: the mock fixtures held code only under app/, so the leak was never seen. Mock tools only. Not a per-incident price.
The staging BigQuery table held rows for the live run: 5 LLM_REQUEST, 5 LLM_RESPONSE, 2 TOOL_COMPLETED. The five model calls match the local run's shape.
“66 `agent_events` rows: 5 `LLM_REQUEST`, 5 `LLM_RESPONSE`, 2 `TOOL_COMPLETED` and others.”
Limits. What the run proposed, its time and the approval were owner-observed, not controller-confirmed. The proposal was needs_more_data and did not cite the leak. Cost was not measured then.
Run 3f60dd3f found the leak at confidence 0.95 and proposed a try/finally patch; with the critic loop, all three live pool-fault runs proposed that patch.
“Run 3f60dd3f: the exit check passed. The model found the leak at confidence 0.95”
“never changed (`code_patch` on `app/checkout_v2.py` with a `try/finally`”
Limits. A handful of runs, not a rate. The first three live runs before 3f60dd3f ended unresolved, and its patch would not apply cleanly to the real file. A later run's patch applied at the deployed commit.
In the spike graph, critique reached follow_up_agent, the reviser and critic_2; the gate received the revised Diagnosis; no save_proposal node was needed.
“the gate receives the revised `Diagnosis`; no `save_proposal` node needed”
Limits. This is the test-only spike topology. The production graph's loop is a separate claim.
Since Phase 1 the graph gained change_agent, history_agent, the four-role critic loop, validate_proposal, report_unresolved, apply_mitigation, verify_outcome, postmortem_agent and store_postmortem_draft.
Limits. Each addition changed the workflow version, so unfinished runs were drained before those deploys. The graph shape is the same for every run; model choice is data.
The critic loop is unrolled: critic_1, then at most one follow-up, one revision and critic_2, then validation and the single approval interrupt, or an unresolved ending.
“Revision rate: 5 of 6. All six reached the gate; none ended unresolved.”
Limits. Six staging runs on one image, a single measurement. The follow-up may make at most three tool calls. Request changes means reject with a comment; an operator can then continue it as a new linked run.
In S1 spike benchmarks on synthetic cases, native Gemini passed specialist (20/20), OpenAI gpt-6-luna passed follow-up (20/20), and Qwen3 on Vertex passed specialist (19/20) and follow-up (19/20).
“PASS: 20/20 valid tool calls, p95 21.5 s (deadline 60 s)”
“`gpt-6-luna` (medium) 20/20”
“`qwen3-235b-a22b-instruct-2507` specialist 19/20 (p95 8.0 s) and `follow_up_agent` 19/20”
Limits. Synthetic cases, one run, no repeats; Qwen3's specialist pass has no margin. Models were called from a local machine; capacity and latency may differ by region and time of day. These are isolated role tests; runs that mix providers are a separate claim. Raw results are in backend/spikes/s1_results.
Specialist tool calling failed for all five OpenAI and OpenRouter models tested; best was 18/20 against a bar of 19. Native Gemini and Qwen3 on Vertex passed.
“none reaches 19/20. Best: `gpt-6-luna` and `qwen3.7-flash` 18/20”
Limits. Same synthetic, one-run caveats. Use as an honesty point, not a model ranking.
gpt-6.1-sol at low and medium effort caught 10/10 wrong diagnoses and passed 3/3 correct ones. Open models caught at most 5/10 on OpenRouter and 8/10 on Vertex.
“`gpt-6.1-sol` at low and medium effort, 10/10 wrong caught, 3/3 correct passed”
Limits. Provisional: the critic was judged on 3 correct cases where the spec asks for 5, and three cases were withdrawn after the run. Open-model critics all failed. Synthetic cases, one run.
Native gemini-3.8-flash was valid first time on 6/6 synthesizer cases at a 6,000-token cap (p95 36.2 s). At the default 3,000 cap it failed one case (5/6).
“6/6 valid first time, p95 36.2 s”
Limits. One run, from a local machine to the Vertex AI global endpoint. The cause of the earlier miss was not retained; thinking headroom is the supported explanation, not a confirmed cause.
Profiles give each role a model chain across vendors, and a failed candidate falls back visibly. A comparison replays one run's frozen evidence under several profiles.
“both `log_agent` turns recorded `SimulatedOutage` at $0 and fell back to Gemini”
“Every column read the source's frozen pool with no gaps.”
Limits. One staging run per profile, not averages. No 429 retry happened on staging; that path is tested locally. Comparison members make no live reads and take no actions.
Specialist, diagnosis, approval-pointer, decision, outcome and other cards render in seq order with evidence chips; no composer is rendered during a run.
“each card is one message with a `data-card` part”
“PASS: no composer is rendered and `onNew` refuses”
Limits. Local (jsdom and Chromium). Evidence chips show IDs only until run_evidence exists. Critic and follow-up cards render from the spike's event shapes, not from a production critic loop. The decision record rows are informational.
Every screen was restyled to the Harbor board and compared at 1440x900 on the local scripted stack. Graph text measured 13.44 px (names) and 12.48 px (mode tags).
“node names 13.44px, mode tags 12.48px”
“Board screens compared (1440×900, full page)”
Limits. The decision record lists states no browser saw (recorded replay, failed run, viewer view, no-role user, API-unreachable sign-in, budget-stopped title). The branch had not been deployed to staging at the time. Desktop only (1440 px); phone widths are not verified.
Each JSON block has a Copy button; 1,066 characters of valid JSON were copied from the synthesizer's output in a browser check.
“Checked in a browser: 1,066 characters of valid JSON were copied from the synthesizer's output.”
Limits. The copy-failure message was not tested.
Operators pick model mode, tool mode and profile, plus an outage flag and small budget for live runs. Start from recorded alert sends a saved alert, labelled, in those modes.
“"Start from recorded alert" made run `dcbd6ada` (scripted, mock), labelled "Recorded alert", and it was left paused.”
“both `log_agent` turns recorded `SimulatedOutage` at $0 and fell back to Gemini”
“`budget_stopped` at the synthesizer: it needed $0.101 with $0.061 left”
Limits. The console has no health widget or fault switch yet; an operator turns faults on from the command line. Cost shows on each run's page, not in the console.
The Approval screen shows the hypothesis board, the critic's review and the before and after of a revision; a comparison shows each profile's result in one view.
“The Approval screen showed the hypothesis board (leading hypothesis, open and ruled-out groups with their evidence rows), the critic's review”
“The comparison was exported to the archive draft with the two published runs and played back read-only from it.”
Limits. Checked on one stored run with no model calls. On the Approval screen the Ask panel is near the page end. The Evals screen is not covered by this claim.
Planned for Phase 5: an audit view with a span timeline, per-node figures and a two-run overlay on the run graph, and a Looker Studio dashboard. Replay export: Phase 6.
Limits. Not built. Spans Phases 5 and 6.
Two sweeps arrived through Cloud Scheduler OIDC in the check window, both answered 200. The sweep is service-wide, not tied to the test run.
“2 sweeps in the check's window, both 200 (service-wide, not tied to the run)”
“"sweeper_requests_in_window": [ 200, 200 ]”
Limits. Shows authentication and routing only. The decision record row is informational.
Two execute requests (start and resume) passed Cloud Tasks OIDC, both 200, longest 43 s. The 900 s dispatch deadline equals Cloud Run's request timeout.
“2 execute requests for the run (start and resume), both 200, the longest 43 s; dispatch deadline 900 s equals Cloud Run's request timeout”
“"max_s": 43.017080195, "dispatch_deadline_s": 900, "cloud_run_timeout_s": 900”
Limits. Scripted models, so 43 s says nothing about real model latency.
Staging proof runs: switch to full pool in 39.0, 38.2 and 36.5 s; local 36.3 and 36.4 s. Then pool-timeout errors and 503s after about 5.0 s.
“Switch to full pool 39.0 s.”
“Switch to full pool 36.5 s”
“36.3 s and 36.4 s from the switch to a full pool”
Limits. The fault is a real request-scoped session leak, switched on by a feature flag. Timings are with the synthetic traffic job. About 37 seconds is the owner's walkthrough figure, not the design's 40-45 s. Alert latency is unmeasured. The records carry no date.
In proof run 1 the health endpoint reported recovered by the first poll (about 17 s) after the pool limit was raised from 10 to 80.
“Mitigation: `recovered` by the first poll (about 17 s)”
Limits. The mitigation was applied through the storefront's admin controls by an operator, not by an approved agent action; the agent-applied case is a separate claim. recovered means a checkout succeeded after the mitigation, within 30 s. The record has no date.
After the mitigation the old pool still held its leaked connections: 10 in the local check; retired_in_use 11 throughout proof run 3 on staging.
“including the old pool still holding 10 after the mitigation”
“`retired_in_use` 11 throughout”
Limits. The 11th connection's cause is read from log timing, not from the database. A residual undercount is possible (from code review, not reproduced).
The health endpoint reported estimated headroom of 318.5 s, 329.0 s and 332.0 s (5.3 to 5.5 min) after the mitigation in staging proof runs 1 to 3.
“headroom 318.5 s (5.3 min)”
“headroom 329.0 s (5.5 min)”
“headroom 332.0 s (5.5 min)”
Limits. This is the health endpoint's estimate; no run waited out the headroom to measure it. A slow demo after the mitigation can run out again.
POST /admin/reset returned generation 1 with nothing checked out and no faults in 17 s. The traffic job takes about 3 minutes from start to first request.
“Reset in 17 s: generation 1, nothing checked out, no faults”
“About 3 minutes from `traffic-start` to the first request”
Limits. This record covers the storefront reset. Closing agent pull requests and rolling back mitigations is the separate environment reset, a command-line target; its admin page is Phase 6.
GitHub Actions CI took 78 s on 53f8e11 and 76 s on 6c4b9d1; storefront tests went from 21 to 23 with the pool-count fix.
“78 s on `53f8e11` (first run); 76 s on `6c4b9d1`; storefront tests 21, then 23 with the pool-count fix”
Limits. Branch protection is off; GitHub refused it on the free plan. That GitHub never holds Google Cloud credentials is a design decision, not a result. The record has no date.
A Monitoring alert started each fault's run with no click; after approval action-runner applied the mitigation, opened a PR through the GitHub App, and CI status arrived by webhook.
“The push arrived and the run was created at 13:35:31Z (log: `started`). It paused at the gate.”
“The UI went from "CI pending" to "CI passed".”
Limits. Two staging runs, one per fault. An alert took about 3.5 to 5 minutes to start a run, almost all of it inside Monitoring. The app never merges a pull request.
Through Cloud Run and the proxy: first byte 406 ms, lag p50 402 ms (max 639), pings every 15 s, close at 240.7 s, clean Last-Event-ID reconnect.
“pings every 14.9–15.1 s while paused, so nothing buffers”
“"lag_ms": { "p50": 402, "max": 639 }”
Limits. Scripted models. Lag measures created_at to arrival through a 0.5 s poll, so it is transport lag, not model latency. The decision record row is informational.
At paused, queried with no wait, the BigQuery rows included HITL_INPUT_REQUEST and INVOCATION_COMPLETED; after a revision change (00005 to 00006) the later events were present.
“at `paused`, queried with no wait, the rows include `HITL_INPUT_REQUEST` and `INVOCATION_COMPLETED`”
“"triage-api-00005-qk7", "triage-api-00006-47c"”
Limits. That flushing happens before the executor writes the status rests on source reading, not on a blocked-flush experiment.
After a fix (the tracer resource now names the project), the rows carried 2 distinct trace ids and the one fetched returned 200 from Cloud Trace with ADK span names.
“the one fetched returned 200 from Cloud Trace with ADK span names”
“"distinct_trace_ids": 2, "trace": { "status": 200,”
Limits. One trace id fetched. Before the fix every span export got 400 from the Telemetry API. The decision record row is informational.
Planned for Phase 5: a Looker Studio dashboard on the BigQuery views.
Limits. Not built.
ADK DatabaseSessionService (asyncpg) on Omni 17.9.0: session and event survive a new service instance and a real alembic upgrade head. ADK tables stay in public.
“PASS (ADK tables stay in `public`, `alembic_version` in `triage`)”
Limits. Local Omni only here; the staging VM runs the same image. Omni's free license is for developing, testing, prototyping and demonstrating. A pilot on client data needs managed AlloyDB or a paid Omni subscription. The record gives no date.
Hybrid query tests passed for stable-only, service filter, taint, active index version and one hit per document; the simple+english index matches CamelCase, stemmed English and snake_case.
“S5 hybrid query: stable-only, service filter, taint, active version, one hit per document”
“PASS (tsv matches CamelCase exact, english stem, snake_case)”
Limits. These are the spike's search tests; history_agent and the knowledge loader came later. The records give no date.
With real gemini-embedding-001 vectors on 12 documents and 10 queries, hybrid search ranked the right document top-3 in 10/10 and top-1 in 9/10. Exact search, no ScaNN.
“hybrid top-3 10/10, top-1 9/10; vector-only top-3 10/10; no ineligible document returned”
Limits. 12 documents is tiny. ScaNN recall with real embeddings stays unmeasured; v1 uses exact search because filtered ScaNN recall was 0.43 against a 0.95 bar. v1 does not use ScaNN. The record gives no date.
First boot 137 s to ready; warm start 25 s (55-61 s from env-up to a healthy database); a snapshot-restored VM answered at 112 s with the newest run completed.
“Omni answered 137 s after kernel start (`omni-report: ready uptime_s=137`)”
“a VM booted from a snapshot of the live data disk answered 112 s into its boot”
Limits. Informational rows with no sample counts. The first snapshot drill attempt failed in the drill script itself (fixed). The formal restore drill is Phase 6. No high availability.
Terraform deployed Phase 1 and the storefront. make env-down left the VM TERMINATED, queue and sweeper PAUSED, health database_unavailable.
“VM `TERMINATED`, queue and sweeper `PAUSED`, health `database_unavailable`”
“`make env-down` cancelled traffic each time and stopped the VM”
Limits. Cost savings are not measured: the cited records contain no billing data. Terraform drift of gcloud client metadata remained after the restart steps. A fresh project needs two applies; fresh-project stand-up is a Phase 6 item.
Over the first session, the traffic job sent 2,317 requests, all completed, 176 returned 5xx (all during the two incidents), 0 failed; make env-down cancelled the job each time.
“2,317 requests sent and completed, 176 5xx (all during the two incidents), 0 failed”
Limits. env-down printed no storefront traffic to stop both times although the execution was cancelled; the cause of the non-zero exit is still unverified. The record has no date.
Planned for Phase 6: restore drill, auto-stop test and host metrics. Omni here is a showcase with no HA, self-patched; managed AlloyDB for PostgreSQL is the recommended production target.
Limits. A recommendation, not an experiment. The code is expected to be unchanged because it uses only shared features; that is not verified.
FIGURES
The numbers quoted on the site, each with where it was recorded, the date where the record gives one, and which line of the decision record states it. They come from a few recorded runs and tests, not a benchmark; each card says which.
From the fault switch to an exhausted connection pool
3 runs
From an operator's mitigation to recovered
measured in proof run 1; applied by an operator; agents apply it in Phase 2
Estimated time before the pool would run out again after the mitigation
estimated by the health endpoint, not measured to re-exhaustion
Traffic sent to the shop
during the first recorded session; the errors were all during the two incidents
Of worker death recovered, without repeating finished steps
a model error, a cancelled task and a killed process
MILESTONES
Every In build claim names its phase. Phases 1 to 4 are done; Phase 5 is in progress.
Done: durable runs, exact approvals, audit, redaction, spend caps.
Done: agents on the real shop, a separate write service, live mitigation and pull requests.
Done: several model providers, with routing, fallback and side-by-side comparison.
Done: memory and runbooks (4a); a critic that challenges the diagnosis (4b).
In progress: injection screening and evals are done; free-text answers, the Ask screen and the audit view are next.
Launch readiness: rehearsals, dashboards and a restore drill.
BACK TO THE START
Each group in the ledger links to the page that explains it. The architecture page is the place to start: the services, the data stores and the path to production.
Back to Architecture →