AGENTIC AI / GOOGLE CLOUD
Investigate application outages with AI.
When an application starts failing, engineers need to find out why. Beagle Incident Triage examines logs, source code, recent changes and past incidents to propose a cause and response. An engineer reviews the proposal before the system makes an approved change.
- AI DECIDES
- AI chooses what evidence to examine next.
- CODE + ENGINEERS DECIDE
- Code and human approval control what can change.

Customers cannot complete checkout.
How this step works
Beagle Incident Triage receives the alert and starts gathering evidence. An alert alone does not authorise a change.
Checkout waits for database connections
Cloud Monitoring detects failed requests
An authenticated alert starts the investigation
Find why database connections are running out.
How this step works
AI agents examine logs, source code, recent changes and reviewed runbooks. They can ask for more evidence and revise the suspected cause. Unresolved objections stop the investigation.
Logs show connection timeouts
An initial idea is to increase capacity
A second agent asks which code fails to release connections
An engineer reviews the proposed change.
How this step works
The engineer reviews the evidence and response before any change. Code checks that execution matches the approved plan; a digest, or content fingerprint, binds that approval.
Temporary relief and how to undo it
A proposed code fix with supporting citations
Risks, approval expiry and the exact plan
Restore checkout, then review the lasting fix.
How this step works
A separate service applies the approved temporary change and records whether service recovers. It never merges or deploys the code fix; engineers review that pull request separately.
Apply the approved, reversible mitigation
Check whether checkout is healthy again
Open a fix pull request for CI and engineering review
Ordered sequence · evidence, review and outcomes remain visible.
Evidence / agent workWaiting for a personFailingRecovered
See the evidence, response and outcome.
Each archived investigation preserves its evidence, model configuration and outcome. This read-only archive works independently of the showcase database VM.
WHAT SUPPORTS AN INVESTIGATION
From an outage alert to a reviewed response.
Six capabilities connect evidence gathering, human control and reliable execution. Each links to the architecture and its tradeoffs.
REAL INTEGRATIONS
Connect alerts to useful evidence
Cloud Monitoring starts an investigation; Cloud Logging and GitHub supply logs, code and changes. Slack reports progress toward a reviewed response.
Trace the integrations →GOVERNED RAG
Find relevant runbooks and past incidents
Retrieval-augmented generation (RAG) searches reviewed troubleshooting guides and past incident reports. Cited passages help engineers check the proposed cause.
See how knowledge becomes evidence →MODEL RESILIENCE
Continue when a model is unavailable
Route investigations across Gemini, OpenAI and open models hosted on Vertex AI. A fallback attempt must respect data permissions and its reserved budget.
Explore model routing →GOVERNANCE
Keep production changes under human control
Roles determine who can approve. Code screens hostile or sensitive content and validates the plan. Only a separate service holds change credentials.
Inspect the control boundaries →EVALUATIONS
Check whether investigations help
Test recorded incident cases and held-out search questions. Compare diagnosis quality, recovery, latency and estimated cost with the model and evidence configuration attached to each result.
Open recorded evidence →DURABLE GCP EXECUTION
Resume work after an interruption
Cloud Tasks and stored investigation state keep work independent of an open browser. Leases, scheduled recovery checks and reconciliation handle interruptions; external calls may repeat.
See how interrupted work resumes →CHOICES, NOT JUST COMPONENTS
Why this architecture?
How do agents find the right runbook?
An exact error name and a description such as “checkout waits” are different clues. Combining keyword and meaning-based search adds complexity that retrieval tests must justify.
Read the search tradeoff →AUTHORITYWhy can’t an AI investigator change production?
The service consulting AI models has no credentials to change production. A separate service executes approved changes, adding an operational boundary to maintain.
See the approval boundary →EXECUTIONWhat happens if the engineer closes the tab?
An investigation continues from stored state. Queued work and recovery checks make this possible, but add storage, retries and reconciliation to operate.
See how work resumes →WHERE THIS PATTERN FITS
Use AI agents when an outage needs investigation.
A timeout may come from a traffic spike, a database problem or code that retains connections. Agents choose follow-up questions as evidence arrives. A known fault may only need a fixed runbook; occasional debugging may suit an engineer with an AI assistant. Copilot also supports agents, tools and MCP. This system adds a repeatable outage workflow with explicit approval and execution controls.