Under the hoodarchitecture
← Architecture showcase

UNDER THE HOOD / ARCHITECTURE

What runs where,
and why.

Beagle Incident Triage runs as four services on Cloud Run around one AlloyDB Omni database. A queue keeps each investigation going without a browser, and a separate service holds the GitHub and shop write credentials. This page shows what a request touches, what the database stores, and what changes in your production.

SERVICES

What does a request touch?

A person signs in on triage-web, which calls triage-api. Cloud Tasks runs each investigation on triage-api, the run state lives in the database, and approved writes go through action-runner.

ILLUSTRATIVE SEQUENCE01 / 07
A person signs in and starts a run. What does the request touch?
enqueueexecute tasksweeperID tokenalert webhookapproved writesbrowserwebCloud Runtriage-apiCloud RunCloud SchedulerscheduleCloud TasksqueueVertex AIGemini, QwenAlloyDB Omnione VMFirebase Authsign-in, rolesstorefrontdemo shopNo model · write credentials onlyaction-runner
ID tokenenqueueexecute tasksweeperalert webhookapproved writesbrowserFirebase Authsign-in, roleswebCloud Runtriage-apiCloud RunCloud TasksqueueCloud SchedulerscheduleAlloyDB Omnione VMVertex AIGemini, Qwenstorefrontdemo shopNo model · write credentials onlyaction-runner
BROWSER · FIREBASE AUTH

Sign in with Google.

Google sign-in. triage-api checks the ID token on every request.

IDENTITY

Google sign-in through Firebase Auth

triage-api checks the ID token on every request

CLOUD RUN · triage-web

Reach the screens people use.

The screens people use: a Next.js app with a thin server layer that keeps the sign-in session and passes calls on to triage-api. It is the client-facing product surface.

triage-web

A Next.js app with a thin server layer

Keeps the sign-in session

Passes calls on to triage-api

CLOUD RUN · triage-api

One place decides what a run may do.

Our own FastAPI app around the ADK App and Runner. It checks sign-in and roles, runs the run lifecycle, streams events and takes approvals. Clients never talk to ADK directly.

triage-api

FastAPI around the ADK App and Runner

Checks sign-in and roles; takes approvals

ADK's own HTTP endpoints are not mounted

CLOUD TASKS

Queue the run, then execute it.

The run queue. It calls triage-api's private execute endpoint, so a run does not depend on an open browser.

THE RUN QUEUE

enqueue: triage-api adds a task

execute task: Cloud Tasks calls triage-api's private endpoint

CLOUD SCHEDULER

Recover runs that lost their executor.

Cloud Scheduler calls the sweeper, which re-enqueues runs that lost their executor.

THE SWEEPER

Cloud Scheduler calls the sweeper

The sweeper re-enqueues runs that lost their executor

ALLOYDB OMNI · VERTEX AI

Keep run state in the database; call the models.

Everything persistent lives in one AlloyDB Omni instance on a Compute Engine VM. Each run's profile picks its models: Gemini and Qwen on Vertex AI, and OpenAI and OpenRouter models.

DATA AND MODELS

AlloyDB Omni on one Compute Engine VM

Vertex AI: Gemini and Qwen; OpenAI and OpenRouter too

ALERTS · action-runner

Alerts start runs; approved writes go through action-runner.

action-runner makes every write: pull requests through a GitHub App, and approved mitigations through storefront-api's admin API. The process that talks to models holds no write credentials.

ALERTS AND WRITES

storefront alert webhook into triage-api

action-runner: no model, write credentials only

Only IAM-authenticated calls from triage-api

Ordered sequence · evidence, review and outcomes remain visible.

Evidence / agent workWaiting for a personFailingRecovered

The whole system: every service and connection in v1. Dashed card and line: not built yet (Slack notifications, Phase 6). The outline groups the services on Cloud Run.
Cloud RunID tokenrelaysweeperenqueue runexecuteDirect VPCegresshealth check,fault switchIAM-auth callapproved mitigation,rollbackwrite_operationsfault and mitigation statenotificationspush, OIDCread logslogs5xx alertPRworkflow_run webhookbrowserNext.js appFirebase AuthCloud SchedulerscheduleCloud Tasksrun queuetriage-webNext.jstriage-apiFastAPI + ADK Runnerstorefront-apishop API · fault switchmitigation admin APIaction-runnerprivate, deterministicVertex AIGemini, open modelsOpenAI APISecret ManagerAlloyDB Omnion a VMADK sessionsvectors · runsaudit · budgetsevalsSlackchannelPhase 6Pub/SubtopicCloud LoggingTrace / MonitoringGitHubincident-triage-storefront repo

triage-web

Built

The screens people use

The screens people use: a Next.js app with a thin server layer that keeps the sign-in session and passes calls on to triage-api.

Why: It is the client-facing product surface.

triage-api

Built

Decides what a run may do

Our own FastAPI app around the ADK App and Runner. It checks sign-in and roles, runs the run lifecycle, streams events, takes approvals, and receives alerts and CI webhooks. Its agents read logs, code, recent changes and the knowledge base.

Why: One place decides what a run may do. ADK's own HTTP endpoints are not mounted, so clients never talk to ADK directly.

action-runner

Built

Makes the GitHub and shop writes

A small private service that makes every write: pull requests through a GitHub App, and approved mitigations through storefront-api's admin API. It owns the write_operations rows. Only IAM-authenticated calls from triage-api and the operator can reach it.

Why: Credential isolation: the process that talks to models holds no GitHub or mitigation write credentials. Each service has its own database role (see Security).

storefront-api

Built

A realistic shop with real faults

A realistic shop API with a fault switch and an admin API for allowlisted, reversible mitigations. It runs as one instance, so the connection-pool figures in its logs are true. Its code lives in its own repo.

Why: It makes real incidents and real logs on cue, and an approved mitigation can really restore service. An operator switches its faults; the agents read its logs and code, and approved mitigations reach it through action-runner.

MANAGED GOOGLE SERVICES

Which Google services does it lean on?

Each service does one job. None of them is the record of an investigation's decisions; that record stays in the database.

Cloud Tasks

The run queue. It calls triage-api's private execute endpoint.

Cloud Scheduler

Calls the sweeper, which re-enqueues runs that lost their executor.

Firebase Auth

Google sign-in. triage-api checks the ID token on every request.

Vertex AI

Gemini and open models such as Qwen. OpenAI and OpenRouter models are called with API keys from Secret Manager.

Secret Manager

Database passwords and API keys. No service-account JSON keys.

BigQuery

Agent analytics rows, with prompt and response content logged after redaction.

DATA

Why one database?

Everything persistent lives in one AlloyDB Omni instance on a Compute Engine VM. Omni is Google's downloadable, PostgreSQL-compatible build of AlloyDB. One transactional store keeps a decision next to the evidence and approvals behind it.

ADK's own session tables, through DatabaseSessionService

Built

Agent sessions

ADK creates and owns them. Our migrations never touch them.

runs (with the executor lease), run_events, approvals, audit_log, budget_ledger, workflow_versions

Built

Runs and their record

run_events is the sanitized projection the UI reads.

storefront_state

Built

Shop state

Active faults and mitigations, read by every storefront instance.

write_operations

Built

Write results

Action idempotency, owned by action-runner.

kb_documents, kb_chunks, knowledge_loads

Built

Knowledge base

Search is exact vector search plus full-text. There is no ScaNN index in v1. Embeddings are made in triage-api and only the vectors are stored.

model_calls

Built

Cost ledger

Model, tokens, latency and an estimated cost per attempt, including fallbacks and repairs.

eval_runs, eval_results, run_questions

Built

Evals and questions

Run questions are kept apart from the run's ADK session.

Omni's free licence covers developing, testing, prototyping and demonstrating. That is why this environment is a demonstration, and why the rule at the end of the next section exists.

FROM THIS DEMO TO YOUR PRODUCTION

What changes in your production?

This demo runs on self-managed AlloyDB Omni. The recommended target is managed AlloyDB for PostgreSQL, the same engine with high availability, backups and IAM built in. It is a target, not a ready deployment.

The same diagram, read for the boundary: action-runner holds the write credentials and calls no model
enqueueexecute tasksweeperID tokenalert webhookapproved writesbrowserwebCloud Runtriage-apiCloud RunCloud SchedulerscheduleCloud TasksqueueVertex AIGemini, QwenAlloyDB Omnione VMFirebase Authsign-in, rolesstorefrontdemo shopNo model · write credentials onlyaction-runner
ID tokenenqueueexecute tasksweeperalert webhookapproved writesbrowserFirebase Authsign-in, roleswebCloud Runtriage-apiCloud RunCloud TasksqueueCloud SchedulerscheduleAlloyDB Omnione VMVertex AIGemini, Qwenstorefrontdemo shopNo model · write credentials onlyaction-runner
Identity and access, data location, operations, backups, availability and cost: this demo against your production
AspectThis demoYour production
Identity and accesstriage-api cannot be called without Google credentials (checked on staging). People sign in with Firebase, and roles live in the database. triage-api and action-runner each log in to the database with their own restricted role, and a separate account runs migrations. The Cloud SQL connector and IAM database login do not exist on Omni.Managed AlloyDB has IAM built in, so each service could use IAM database login instead of a password.
Where data goesAll state is in one Omni database on a VM with a private IP only. Cloud Run reaches it through Direct VPC egress, and a firewall allows port 5432 only from those subnets. Incident text is redacted before it reaches ADK. Tool output is redacted before a model sees it. Model calls go to Vertex AI, OpenAI and OpenRouter, each per the run's profile. Analytics rows go to BigQuery with redacted content.The same layout on managed AlloyDB. For data-residency needs, Omni also runs on-premises or on another cloud.
Who operates whatWe do. We start the VM for client sessions and stop it afterwards, and we patch the OS and Omni ourselves.Managed AlloyDB for PostgreSQL: the same engine, with high availability, backups and IAM built in.
BackupsThe design is scheduled persistent-disk snapshots, plus a pg_dump before each prod deploy. A cold boot and a restore from snapshot were observed on staging, without sample counts. The snapshot schedule and a formal restore drill are Phase 6.Built into managed AlloyDB.
AvailabilityNo high availability. That is acceptable for client sessions and rehearsals. While the VM is stopped the whole app is offline by design and the UI shows "Environment offline". Paused runs survive a stop and start because the disk persists.Managed AlloyDB with high availability. This is a recommendation, not an experiment. The code is expected to be unchanged because it uses only features both editions share; that is not verified.
Model costEstimated per attempt from a versioned price table; on staging, about $0.03 to $0.16 per runThe same estimate, reconciled against the billing export and provider invoices.
Infrastructure costVM, Cloud Run and Cloud Tasks; stopped between sessionsNot estimated here.

The rule for real data

A pilot that processes a client's real data is not "demonstrating". Real client data needs managed AlloyDB or a paid Omni subscription.

Production target: managed AlloyDB for PostgreSQL.

NEXT

Follow one investigation through the agents.

The agent workflow page shows which agents run, what a person approves, and which model answers each role.

Next: AI investigation →