AGENTIC AI / GOOGLE CLOUD

Investigate application outages with AI.

When an application starts failing, engineers need to find out why. Beagle Incident Triage examines logs, source code, recent changes and past incidents to propose a cause and response. An engineer reviews the proposal before the system makes an approved change.

AI DECIDES
AI chooses what evidence to examine next.
CODE + ENGINEERS DECIDE
Code and human approval control what can change.
The Beagle mascot
ILLUSTRATIVE SEQUENCE01 / 04
An online store’s checkout starts failing. What happens next?
AlertEvidenceAI investigatesEngineer approvesRecovercheckoutstorefront-apirequests failingalert closedlogsCloud Loggingsource codeGitHubrecent changesGitHubrunbooksAlloyDBAI investigationtriage-api · ADKcause: connection leakengineer approvaltriage-webevidence · risk · rollbackapprovedNo changes before this gateapproved reliefaction-runnercheckout healthyfix pull requestGitHubCI + separate reviewNo automatic merge or deploy
AlertEvidenceAI investigatesEngineer approvesRecovercheckoutstorefront-apirequests failingalert closedlogsCloud Loggingsource codeGitHubrecent changesGitHubrunbooksAlloyDBAI investigationtriage-api · ADKcause: connection leakengineer approvaltriage-webevidence · risk · rollbackapprovedNo changes before this gateapproved reliefaction-runnercheckout healthyfix pull requestGitHubCI + separate reviewNo automatic merge or deploy
MONITORING · CLOUD MONITORING + PUB/SUB

Customers cannot complete checkout.

How this step works

Beagle Incident Triage receives the alert and starts gathering evidence. An alert alone does not authorise a change.

EXAMPLE · ONLINE STORE

Checkout waits for database connections

Cloud Monitoring detects failed requests

An authenticated alert starts the investigation

AI INVESTIGATORS · GOOGLE ADK

Find why database connections are running out.

How this step works

AI agents examine logs, source code, recent changes and reviewed runbooks. They can ask for more evidence and revise the suspected cause. Unresolved objections stop the investigation.

FROM CLUES TO A PROPOSED RESPONSE

Logs show connection timeouts

An initial idea is to increase capacity

A second agent asks which code fails to release connections

ON-CALL ENGINEER · HUMAN REVIEW

An engineer reviews the proposed change.

How this step works

The engineer reviews the evidence and response before any change. Code checks that execution matches the approved plan; a digest, or content fingerprint, binds that approval.

WHAT THE ENGINEER SEES

Temporary relief and how to undo it

A proposed code fix with supporting citations

Risks, approval expiry and the exact plan

SEPARATE EXECUTION SERVICE · HEALTH CHECK

Restore checkout, then review the lasting fix.

How this step works

A separate service applies the approved temporary change and records whether service recovers. It never merges or deploys the code fix; engineers review that pull request separately.

RELIEF NOW · CODE REVIEW NEXT

Apply the approved, reversible mitigation

Check whether checkout is healthy again

Open a fix pull request for CI and engineering review

Ordered sequence · evidence, review and outcomes remain visible.

Evidence / agent workWaiting for a personFailingRecovered

RECORDED INVESTIGATIONS

See the evidence, response and outcome.

Each archived investigation preserves its evidence, model configuration and outcome. This read-only archive works independently of the showcase database VM.

View recorded investigations

WHAT SUPPORTS AN INVESTIGATION

From an outage alert to a reviewed response.

Six capabilities connect evidence gathering, human control and reliable execution. Each links to the architecture and its tradeoffs.

01
alertfiringagentsaskinglogscodeFollow real evidence
alertfiringagentsaskinglogscodeFollow real evidence

REAL INTEGRATIONS

Connect alerts to useful evidence

Cloud Monitoring starts an investigation; Cloud Logging and GitHub supply logs, code and changes. Slack reports progress toward a reviewed response.

Trace the integrations →
02
error namesymptomsRRFrunbookcitedOne cited troubleshooting guide
error namesymptomsRRFrunbookcitedOne cited troubleshooting guide

GOVERNED RAG

Find relevant runbooks and past incidents

Retrieval-augmented generation (RAG) searches reviewed troubleshooting guides and past incident reports. Cited passages help engineers check the proposed cause.

See how knowledge becomes evidence →
03
fallbackmodel Atryingunavailablemodel BansweredCheck permission and budget
fallbackmodel Atryingunavailablemodel BansweredCheck permission and budget

MODEL RESILIENCE

Continue when a model is unavailable

Route investigations across Gemini, OpenAI and open models hosted on Vertex AI. A fallback attempt must respect data permissions and its reserved budget.

Explore model routing →
04
AI planproposedapprovalwaitingapprovedactionappliedHuman approval is the gate
AI planproposedapprovalwaitingapprovedactionappliedHuman approval is the gate

GOVERNANCE

Keep production changes under human control

Roles determine who can approve. Code screens hostile or sensitive content and validates the plan. Only a separate service holds change credentials.

Inspect the control boundaries →
05
answer vs outcomecase-01matchescase-02partlycase-03matchesCompare quality and cost
answer vs outcomecase-01matchescase-02partlycase-03matchesCompare quality and cost

EVALUATIONS

Check whether investigations help

Test recorded incident cases and held-out search questions. Compare diagnosis quality, recovery, latency and estimated cost with the model and evidence configuration attached to each result.

Open recorded evidence →
06
queuetaskstatesavedworkerinterruptedresumedRecover interrupted work
queuetaskstatesavedworkerinterruptedresumedRecover interrupted work

DURABLE GCP EXECUTION

Resume work after an interruption

Cloud Tasks and stored investigation state keep work independent of an open browser. Leases, scheduled recovery checks and reconciliation handle interruptions; external calls may repeat.

See how interrupted work resumes →

CHOICES, NOT JUST COMPONENTS

Why this architecture?

Compare architecture alternatives →

WHERE THIS PATTERN FITS

Use AI agents when an outage needs investigation.

A timeout may come from a traffic spike, a database problem or code that retains connections. Agents choose follow-up questions as evidence arrives. A known fault may only need a fixed runbook; occasional debugging may suit an engineer with an AI assistant. Copilot also supports agents, tools and MCP. This system adds a repeatable outage workflow with explicit approval and execution controls.