viewer
The lowest role. It can look at runs. It cannot start a run, decide or administer.
UNDER THE HOOD / SECURITY AND APPROVALS
People sign in with Google and act within their role. An approval is bound to the exact proposal. The service that talks to AI models holds no credential that can change GitHub or the shop; a separate service makes every approved write.
SIGN-IN AND ROLES
People sign in with Google through Firebase Auth. triage-api verifies the ID token on every request. Roles live in the database. The UI hides controls by role, and the API enforces them.
The lowest role. It can look at runs. It cannot start a run, decide or administer.
Can start runs and trigger faults.
The console has no fault switch yet; an operator turns faults on from the command line.
Can approve action plans. Can also verify, deprecate, re-verify or clear taint on knowledge.
Everything above, plus user management and the environment reset.
User management is built. The environment reset is a command-line target; its admin page is Phase 6.
Only Google sign-ins with a verified email are accepted. Email-link sign-in for guests without Google accounts is deferred.
APPROVALS
An approval counts for the exact proposal it was given for. If the proposal changes, the approval no longer applies. How the gate sits in the graph is on the Agent workflow page.
Built
Built
THE WRITE SERVICE
GitHub and shop writes happen outside the process that talks to models. The AI-facing API holds no write credential for GitHub or the shop.
action-runner holds the GitHub App key and the mitigation client, and makes every write. Cloud Run IAM lets only triage-api and the operator call it, and it checks the caller's email per route.
Each service has its own database role. triage-api cannot write action-runner's write_operations table, and action-runner can only read approvals and runs.
Models get four read-only tools. They propose catalog actions with bounded parameters and find-and-replace edits to allowed files; code checks both before the gate and again before writing.
The app opens pull requests but never merges them.
The limit: triage-api still writes approval rows, so a compromised triage-api could forge an approval. The boundary is that it holds no write credential, and action-runner checks the stored approval against the exact proposal.
AUDIT AND REDACTION
Every run leaves a trail. We describe the log as append-only for the application, and we claim nothing stronger.
An administrator can still disable the trigger. Successful writes are recorded in action-runner's own table, not as audit rows. An audit view with a span timeline is planned for Phase 5.
Deterministic patterns remove secrets and personal data. The API applies them to incident text, operator notes and decision comments before any of it reaches ADK. The run's stored events are redacted too.
Four planted secret types, each of which appeared in none of the stored run data or model requests:
Not scanned: the budget ledger and ADK's other tables. Only these four types were planted. A provider error containing password=... is also stored redacted; one pattern was tested.
Every fetched log line, code file and commit is redacted before it is trimmed and before any model sees it. With that in place, prompt and response content is logged to BigQuery on staging.
Prompt-injection screening of the same tool output is a separate control; see the OWASP table below.
BUDGETS
Every model call reserves from the run's budget before it is sent, and that includes retries and fallbacks. If the run cannot afford the call, it is refused and the run ends budget_stopped with no model call.
Ceilings persist across a pause and resume. Cloud Billing budgets stay only as an alert-level backstop.
Each attempt's cost is estimated from a versioned price table and stored, so a run's spend is measured. On staging, each run's summed cost matched its budget ledger. It is an estimate, not the provider's invoice.
OWASP AGENTIC TOP 10
How the controls above map to the OWASP Top 10 for Agentic Applications 2026. Addressed means the main routes into this app are covered and tested; anything listed as still open is bounded by another control. Partly means a known route stays open that no control bounds, and the last column says which.
| Risk | Coverage | How Beagle addresses it | Still open |
|---|---|---|---|
| ASI01 Agent Goal Hijack | Partly | Tool output and incident text are screened for injection, and flagged items are labelled as data. A proposal that cites flagged evidence as support fails validation. The graph is fixed, and a person approves every action.Details
| The screen is five regular expressions; a paraphrase gets past it. Specialists' findings are not screened again. |
| ASI02 Tool Misuse and Exploitation | Addressed | Models get four read-only tools with validated arguments. Mitigations are catalog actions with bounded parameters. Every model attempt reserves budget first.Details
| The specialists' call limits are prompt instructions; the deadline and budget bound them. |
| ASI03 Identity and Privilege Abuse | Partly | Only action-runner holds GitHub and shop write credentials, and each service has its own database role. People sign in and hold roles. Approvals are bound to the exact proposal and re-checked.Details
| Agents share triage-api's identity. The person who started a run can also approve it. |
| ASI04 Agentic Supply Chain Vulnerabilities | Partly | Dependencies install from hash-locked lockfiles. No tools or MCP servers are loaded at runtime. The model registry is checked at startup, and knowledge files are admitted one by one.Details
| No dependency or secret scanning in CI, and no SBOM. Actions and the base image are not pinned by digest. |
| ASI05 Unexpected Code Execution (RCE) | Partly | Models have no code-execution tool in Beagle. They propose bounded find-and-replace edits to allowed files, and changes ship as pull requests that the app never merges.Details
| Model-written code runs in the storefront's pull-request CI before anyone reviews it; that CI's isolation is not assessed here. |
| ASI06 Memory and Context Poisoning | Addressed | Knowledge is screened on load. Search returns only stable, untainted documents: bundle documents are reviewed in Git, and a run's postmortem is searchable only after an approver verifies it. Each run pins its index version.Details
| The same regex screen as ASI01. |
| ASI07 Insecure Inter-Agent Communication | Partly | All agents run in one process and one fixed graph. Every call between services carries a Google ID token checked per route, and GitHub webhooks are signed.Details
| Specialists hand their findings to the synthesizer as free text, with no schema and no second screen. An instruction carried in a finding could sway the proposal. |
| ASI08 Cascading Failures | Addressed | Atomic budgets, an attempt cap, per-specialist deadlines and at most one revision round. An executor that lost its lease cannot finish or pause the run.Details
| Spend is capped per run, not per day; Cloud Billing budgets only alert. |
| ASI09 Human-Agent Trust Exploitation | Partly | The gate shows the checked proposal: its checks, cited evidence, diff, critic objections, confidence and any flags. Citations must be evidence issued to the run. Approvals expire.Details
| Checks show a proposal is well formed and cited, not that it is right. No separation of duties. |
| ASI10 Rogue Agents | Partly | No agent is long-lived or self-starting; every run is bounded. The audit log is append-only for the application. Golden evals test the controls.Details
| A running run cannot be cancelled. No alerts on agent behaviour. |
Checked against the code on 2026-10-06. No third-party governance toolkit is used.
NEXT
The architecture alternatives page sets three Google Cloud paths side by side: ADK on Agent Runtime, ADK on Cloud Run and LangGraph on Cloud Run. Each cell says who carries a production need, with a link to its source.
Next: Architecture alternatives →