Under the hoodGoverned RAG
← Architecture showcase

CAPABILITY 02 / GOVERNED RAG

Find reviewed advice
for an application outage.

Beagle Incident Triage investigates application outages. When an online store’s checkout fails, logs show what happened; a runbook explains what engineers have learned to do about it. Retrieval-augmented generation (RAG) finds relevant runbooks and past incident reports, then gives their cited passages to AI investigators. This design uses AlloyDB to store reviewed knowledge and preserve the evidence behind a proposed response.

Engineers review the adviceSearch by words and meaningSave the cited passage

EXAMPLE · CHECKOUT CANNOT GET DATABASE CONNECTIONS

How does a troubleshooting guide help diagnose an outage?

Follow a runbook: an engineer-written guide to diagnosing and recovering from connection exhaustion. This is an illustrative example, not a recorded investigation or measured result.

ILLUSTRATIVE SEQUENCE01 / 06
From a connection-exhaustion guide to advice an engineer can check

EXAMPLE GUIDE: runbooks/pool-exhaustion.md · checkout waits for database connections

1 · Screen2 · Review3 · Index4 · Retrieve5 · Cite6 · Learna reviewed report joins the indexloaderunsafe → excludedengineer reviewdrafthuman verifiedAlloyDBwords + meaningsearchcheckout timeoutrun_evidenceexact passagepostmortemcause + outcome
1 · Screen2 · Review3 · Index4 · Retrieve5 · Cite6 · Learnloaderunsafe → excludedengineer reviewdrafthuman verifiedAlloyDBwords + meaningsearchcheckout timeoutrun_evidenceexact passagepostmortemcause + outcome
DOCUMENT LOADER · CODE CHECKS

Check a troubleshooting guide before agents can use it.

The loader checks the document’s format and metadata, screens for secrets and instructions aimed at manipulating agents, and keeps unreviewed or tainted content out of search. Passing a format check does not make advice trustworthy.

EXAMPLE · CONNECTION EXHAUSTION RUNBOOK

A runbook explains symptoms and recovery steps

Required metadata, sources and content fingerprint

Invalid documents are quarantined; unsafe content is excluded

ENGINEER · HUMAN KNOWLEDGE REVIEW

An engineer verifies the advice and its source.

Seed runbooks are reviewed in Git. Newly generated incident reports use an audited endpoint restricted to approvers. Each review records a digest, or content fingerprint, so an agent cannot approve its own advice or reuse approval after an edit.

WHAT MAKES THE GUIDE SEARCHABLE

Source: reviewed Git troubleshooting guide

Status: stable · human verified · untainted

Changing the content requires a new review

DOCUMENT LOADER · ALLOYDB + VERTEX AI

Make the reviewed guide searchable by words and meaning.

The loader splits the runbook at headings while retaining title and section context. Vertex AI creates embeddings: numeric vectors representing a passage’s meaning. AlloyDB stores those vectors beside full-text search data. The application calls Vertex AI under its service identity, without a database service-account key.

ALLOYDB · DOCUMENTS AND SEARCH PASSAGES

kb_documents: full text, source and review status

kb_chunks: passages, meaning vectors and word tokens

knowledge_loads: index version, model and content hashes

KNOWLEDGE AGENT · SEARCH TOOL

Find advice for checkout connection timeouts.

Eligibility checks limit both searches to stable, untainted documents in the active index version, with service and corpus filters. One ranking matches words; another matches meaning. Reciprocal rank fusion (RRF) combines their positions rather than comparing incompatible raw scores.

TWO CLUES FROM THE CHECKOUT OUTAGE

Exact error name: DatabaseConnectionTimeout

Symptom: checkout waits for free database connections

Exact vector + full-text rankings → reciprocal rank fusion

EVIDENCE STORE · CODE VALIDATION

Save the exact advice supporting the proposed response.

The passage is treated as evidence data, not instructions to the agent. Code stores it with its source and search metadata in run_evidence for this investigation. Proposal checks reject citations not issued to that investigation. Later index edits cannot alter the saved excerpt.

ILLUSTRATIVE CITATION · CONNECTION EXHAUSTION

kb:runbooks/pool-exhaustion.md#remedy

“Raising capacity buys time; inspect connection release.”

Git runbook · human reviewed · current (illustrative)

Matched by error name and symptom · index version recorded

INCIDENT REPORT AGENT · ENGINEER REVIEW

Turn the checkout outcome into reviewed advice for next time.

A postmortem records what caused the outage, what was attempted and what happened. Its draft stays out of search. After screening and a digest-bound human review, it becomes stable and joins the active index. Past advice supports the next investigation without overriding fresh evidence.

POSTMORTEM · A REPORT OF CAUSE AND OUTCOME

AI drafts a report from saved investigation results

Code adds observed recovery and failed remedies

An engineer reviews it before search can return it

Ordered sequence · evidence, review and outcomes remain visible.

Evidence / agent workWaiting for a personFailingRecovered

The knowledge agent compares past incidents with the checkout outage, including remedies that failed. It can ask for more evidence, but cannot mark advice as trusted, change search eligibility or approve production changes.

UPDATING THE SEARCHABLE KNOWLEDGE

Keep the existing knowledge available while an update is built.

A newly reviewed incident report can join the current search index immediately. Reloading the full collection of runbooks builds a candidate index version first. It carries forward reports created by investigations unless replaced, and removes guides no longer in the collection.

Each load records document hashes, the source Git commit, the passage-splitting version and the embedding model. Changing the model requires recalculating meaning vectors for the whole collection.

What happens if a knowledge update fails?
  1. 01 / SERVINGExisting knowledge stays searchable

    AI investigators continue using the current index of reviewed guides and reports.

  2. 02 / BUILDINGBuild and check the replacement index

    Check document metadata, safety screening, passage boundaries and meaning vectors before switching.

  3. 03 / COMMITSuccess → switch to the checked version

    Failure → keep the previous version active and raise an alert. The active-version pointer changes atomically only after validation.

WHY COMBINE TWO SEARCH METHODS?

Find an exact error name
or a description of the symptom.

Documented v1 choice
KEYWORD PATH

Match a specific error name.

DatabaseConnectionTimeout

An exception name can identify the right guide even when its wording differs from the agent’s question. PostgreSQL full-text search uses simple and english tokens; code extracts error and configuration identifiers and OR-joins them with the search text.

SEMANTIC PATH

Match a similar failure description.

“Checkout requests wait for free connections”

A guide about connection exhaustion may help even if it uses different words. Exact nearest-neighbour search ranks eligible meaning vectors; it can return matches when keyword search finds none. A query vector is reused only for the same query and model within one investigation.

REVIEWED, UNTAINTED CONTENT ONLYWord matches + meaning matches
↓
Reciprocal rank fusion (RRF)

Combine ranking positions.
Keep the best passage from each document.

Alternative: one search method

Using only word matching or meaning matching simplifies operation. Words can miss paraphrased symptoms; meaning can underweight an exact exception name or configuration key.

Why combine them?

RRF combines the two ranking positions without comparing incompatible scores. Both searches filter for approved content first. A bounded penalty lowers the rank of older advice while keeping its need for re-review visible.

What this costs

The application must maintain two searches, passage and embedding versions, and the fusion query. This is the documented design rationale; it is not evidence that hybrid search outperforms either method.

When to reconsider

Simplify if held-out retrieval tests show one method adds little. Test a reranker if useful guides are found but appear too low in the results. Revisit approximate vector search if exact search becomes too expensive at the actual scale and filters.

V1: exact search. No ScaNN index. No reranker.

Approximate vector search can miss useful passages after trust and service filters are applied. V1 uses exact search as the baseline. Any later optimisation must preserve error-name cases and filtered recall: the share of relevant eligible documents found. No Omni benchmark results are published here.

TEST WHETHER SEARCH FINDS THE RIGHT ADVICE

Test the retrieved guides before judging the AI’s answer.

Hold the knowledge Git commit and embedding model fixed. Compare word-only, meaning-only and hybrid search. Measure relevant documents found in the top five (recall@5), the share of those results that are relevant (precision@5), and how highly the first useful match ranks (mean reciprocal rank). Test incidents must be held out from their generated reports.

Database tests can use precomputed query vectors without calling a model. Replaying an AI investigation over recorded evidence tests reasoning, but does not prove the index or eligibility filters work.

What retrieval tests need to catch

  • Exact exception names and paraphrased symptoms
  • Unreviewed drafts, deprecated advice and tainted content must be excluded
  • Older eligible advice must show “needs re-review”
  • Return advice for the correct service, including passages beyond a guide’s first section
  • Distinguish misleading near-matches and remedies that previously failed
  • Reject invented citations and preserve the exact saved excerpt
Inspect recorded evaluation evidence →

WHY ALLOYDB HERE?

Store knowledge, approvals and investigation evidence together.

AlloyDB is PostgreSQL-compatible. It stores investigation sessions, approvals, review metadata, meaning vectors, full-text indexes and saved evidence excerpts. Keeping them in one transactional database helps connect a decision to the knowledge it used. A managed retrieval service remains an alternative behind the search interface, but adds a consistency boundary between systems.

The showcase uses AlloyDB Omni, the self-managed edition, on one Compute Engine VM without high availability. Managed AlloyDB is the production target; migration and production availability need their own validation.

See the wider data architecture →