triage-web
Built
The screens people use
The screens people use: a Next.js app with a thin server layer that keeps the sign-in session and passes calls on to triage-api.
Why: It is the client-facing product surface.
UNDER THE HOOD / ARCHITECTURE
Beagle Incident Triage runs as four services on Cloud Run around one AlloyDB Omni database. A queue keeps each investigation going without a browser, and a separate service holds the GitHub and shop write credentials. This page shows what a request touches, what the database stores, and what changes in your production.
SERVICES
A person signs in on triage-web, which calls triage-api. Cloud Tasks runs each investigation on triage-api, the run state lives in the database, and approved writes go through action-runner.
Google sign-in. triage-api checks the ID token on every request.
Google sign-in through Firebase Auth
triage-api checks the ID token on every request
The screens people use: a Next.js app with a thin server layer that keeps the sign-in session and passes calls on to triage-api. It is the client-facing product surface.
A Next.js app with a thin server layer
Keeps the sign-in session
Passes calls on to triage-api
Our own FastAPI app around the ADK App and Runner. It checks sign-in and roles, runs the run lifecycle, streams events and takes approvals. Clients never talk to ADK directly.
FastAPI around the ADK App and Runner
Checks sign-in and roles; takes approvals
ADK's own HTTP endpoints are not mounted
The run queue. It calls triage-api's private execute endpoint, so a run does not depend on an open browser.
enqueue: triage-api adds a task
execute task: Cloud Tasks calls triage-api's private endpoint
Cloud Scheduler calls the sweeper, which re-enqueues runs that lost their executor.
Cloud Scheduler calls the sweeper
The sweeper re-enqueues runs that lost their executor
Everything persistent lives in one AlloyDB Omni instance on a Compute Engine VM. Each run's profile picks its models: Gemini and Qwen on Vertex AI, and OpenAI and OpenRouter models.
AlloyDB Omni on one Compute Engine VM
Vertex AI: Gemini and Qwen; OpenAI and OpenRouter too
action-runner makes every write: pull requests through a GitHub App, and approved mitigations through storefront-api's admin API. The process that talks to models holds no write credentials.
storefront alert webhook into triage-api
action-runner: no model, write credentials only
Only IAM-authenticated calls from triage-api
Ordered sequence · evidence, review and outcomes remain visible.
Evidence / agent workWaiting for a personFailingRecovered
triage-web
Built
The screens people use: a Next.js app with a thin server layer that keeps the sign-in session and passes calls on to triage-api.
Why: It is the client-facing product surface.
triage-api
Built
Our own FastAPI app around the ADK App and Runner. It checks sign-in and roles, runs the run lifecycle, streams events, takes approvals, and receives alerts and CI webhooks. Its agents read logs, code, recent changes and the knowledge base.
Why: One place decides what a run may do. ADK's own HTTP endpoints are not mounted, so clients never talk to ADK directly.
action-runner
Built
A small private service that makes every write: pull requests through a GitHub App, and approved mitigations through storefront-api's admin API. It owns the write_operations rows. Only IAM-authenticated calls from triage-api and the operator can reach it.
Why: Credential isolation: the process that talks to models holds no GitHub or mitigation write credentials. Each service has its own database role (see Security).
storefront-api
Built
A realistic shop API with a fault switch and an admin API for allowlisted, reversible mitigations. It runs as one instance, so the connection-pool figures in its logs are true. Its code lives in its own repo.
Why: It makes real incidents and real logs on cue, and an approved mitigation can really restore service. An operator switches its faults; the agents read its logs and code, and approved mitigations reach it through action-runner.
MANAGED GOOGLE SERVICES
Each service does one job. None of them is the record of an investigation's decisions; that record stays in the database.
The run queue. It calls triage-api's private execute endpoint.
Calls the sweeper, which re-enqueues runs that lost their executor.
Google sign-in. triage-api checks the ID token on every request.
Gemini and open models such as Qwen. OpenAI and OpenRouter models are called with API keys from Secret Manager.
Database passwords and API keys. No service-account JSON keys.
Agent analytics rows, with prompt and response content logged after redaction.
DATA
Everything persistent lives in one AlloyDB Omni instance on a Compute Engine VM. Omni is Google's downloadable, PostgreSQL-compatible build of AlloyDB. One transactional store keeps a decision next to the evidence and approvals behind it.
ADK's own session tables, through DatabaseSessionService
Built
ADK creates and owns them. Our migrations never touch them.
runs (with the executor lease), run_events, approvals, audit_log, budget_ledger, workflow_versions
Built
run_events is the sanitized projection the UI reads.
storefront_state
Built
Active faults and mitigations, read by every storefront instance.
write_operations
Built
Action idempotency, owned by action-runner.
kb_documents, kb_chunks, knowledge_loads
Built
Search is exact vector search plus full-text. There is no ScaNN index in v1. Embeddings are made in triage-api and only the vectors are stored.
model_calls
Built
Model, tokens, latency and an estimated cost per attempt, including fallbacks and repairs.
eval_runs, eval_results, run_questions
Built
Run questions are kept apart from the run's ADK session.
Omni's free licence covers developing, testing, prototyping and demonstrating. That is why this environment is a demonstration, and why the rule at the end of the next section exists.
FROM THIS DEMO TO YOUR PRODUCTION
This demo runs on self-managed AlloyDB Omni. The recommended target is managed AlloyDB for PostgreSQL, the same engine with high availability, backups and IAM built in. It is a target, not a ready deployment.
| Aspect | This demo | Your production |
|---|---|---|
| Identity and access | triage-api cannot be called without Google credentials (checked on staging). People sign in with Firebase, and roles live in the database. triage-api and action-runner each log in to the database with their own restricted role, and a separate account runs migrations. The Cloud SQL connector and IAM database login do not exist on Omni. | Managed AlloyDB has IAM built in, so each service could use IAM database login instead of a password. |
| Where data goes | All state is in one Omni database on a VM with a private IP only. Cloud Run reaches it through Direct VPC egress, and a firewall allows port 5432 only from those subnets. Incident text is redacted before it reaches ADK. Tool output is redacted before a model sees it. Model calls go to Vertex AI, OpenAI and OpenRouter, each per the run's profile. Analytics rows go to BigQuery with redacted content. | The same layout on managed AlloyDB. For data-residency needs, Omni also runs on-premises or on another cloud. |
| Who operates what | We do. We start the VM for client sessions and stop it afterwards, and we patch the OS and Omni ourselves. | Managed AlloyDB for PostgreSQL: the same engine, with high availability, backups and IAM built in. |
| Backups | The design is scheduled persistent-disk snapshots, plus a pg_dump before each prod deploy. A cold boot and a restore from snapshot were observed on staging, without sample counts. The snapshot schedule and a formal restore drill are Phase 6. | Built into managed AlloyDB. |
| Availability | No high availability. That is acceptable for client sessions and rehearsals. While the VM is stopped the whole app is offline by design and the UI shows "Environment offline". Paused runs survive a stop and start because the disk persists. | Managed AlloyDB with high availability. This is a recommendation, not an experiment. The code is expected to be unchanged because it uses only features both editions share; that is not verified. |
| Model cost | Estimated per attempt from a versioned price table; on staging, about $0.03 to $0.16 per run | The same estimate, reconciled against the billing export and provider invoices. |
| Infrastructure cost | VM, Cloud Run and Cloud Tasks; stopped between sessions | Not estimated here. |
A pilot that processes a client's real data is not "demonstrating". Real client data needs managed AlloyDB or a paid Omni subscription.
Production target: managed AlloyDB for PostgreSQL.
NEXT
The agent workflow page shows which agents run, what a person approves, and which model answers each role.
Next: AI investigation →