Under the hoodcomparison
← Architecture showcase

UNDER THE HOOD / ARCHITECTURE ALTERNATIVES

Where would you
build this?

Who carries each production need, on each framework's recommended Google Cloud path. Sources are checked on the date shown; nothing is scored.

THE THREE COLUMNS

What is being compared?

Three ways to run an agent like this one on Google Cloud. Each is a column in the matrix below.

COLUMN 1

ADK on Agent Runtime

ADK 2.x on Agent Engine, now named Agent Runtime.

COLUMN 2

ADK on Cloud Run

ADK 2.x on Cloud Run, with a database session service.

COLUMN 3

LangGraph on Cloud Run

Open-source LangGraph on Cloud Run, with a Postgres checkpointer.

The LangGraph column is open-source LangGraph on Cloud Run. LangChain's managed platform (LangSmith Cloud) is a different deployment, so it appears only as a row note in the matrix.

THE MATRIX

Who carries each need?

Every cell is what the vendor's own documentation says, with a link to the page and the date it was checked. Each row's notes link to this build's claims.

  • Platformthe managed platform carries it
  • Frameworkthe framework carries it
  • Partialthe documentation covers part of it
  • Your teamyour team carries it
Who carries each production need, per baseline, with this build's status
NeedADK on Agent RuntimeADK on Cloud RunLangGraph on Cloud RunThis build
01A run outlives the request and the instance.
Partial

Long-running query jobs run asynchronously for up to seven days; StreamQuery requests are capped at 10 minutes. Mid-job instance loss is not described.

Source · checked 2026-10-02
Partial

Request-based billing allocates CPU during request processing; instance-based billing allocates it for the instance lifecycle. Idle instances can be shut down at any time.

Source · checked 2026-10-02
Partial

Request-based billing allocates CPU during request processing; instance-based billing allocates it for the instance lifecycle. Idle instances can be shut down at any time.

Source · checked 2026-10-02
3 Proven

If the worker dies mid-run, the run resumes and skips finished steps.

Row notes (3) and this build's claims

Cloud Run: a request timeout can be extended up to 60 minutes; the default is 5 minutes. Source · checked 2026-10-02

Cloud Run: jobs run tasks to completion and do not serve requests; task timeout extends up to 168 hours. Source · checked 2026-10-02

LangSmith Cloud: fully managed infrastructure for stateful, long-running agents with persistent state and background execution. Source · checked 2026-10-02

The LangGraph column is open-source LangGraph on Cloud Run. LangChain's managed platform (the LangSmith Cloud note above) is a different deployment, so it appears only here, as a row note.

02A pause for human approval lasting hours or days.
Partial

RequestInput pauses a node and the reply becomes its output. The page does not state how long a pause can last or its storage needs.

Source · checked 2026-10-02
Partial

RequestInput pauses a node and the reply becomes its output. The page does not state how long a pause can last or its storage needs.

Source · checked 2026-10-02
Framework

interrupt() saves state through a checkpointer and the graph waits indefinitely until resumed with Command(resume=...).

Source · checked 2026-10-02
3 Proven

A paused run resumes correctly after the service restarts, in tests.

Row notes (2) and this build's claims

Agent Runtime: Agent Platform Sessions keep events for a configurable TTL, with a default of 365 days. Source · checked 2026-10-02

ADK: DatabaseSessionService stores session data in a database, and the data survives application restarts. Source · checked 2026-10-02

03Recovery when a worker dies mid-run.
Partial

ADK provides resume by invocation id; Agent Runtime pages don't describe it. Detecting the failure and triggering the resume are not described.

Source · checked 2026-10-02
Partial

ADK provides resume by invocation id with ResumabilityConfig. Detecting the failure and triggering the resume are not described.

Source · checked 2026-10-02
Partial

Sync durability writes each checkpoint before the next step; async risks an unwritten checkpoint on a crash. Detecting a crash is not described.

Source · checked 2026-10-02
4 Proven

If the worker dies mid-run, the run resumes and skips finished steps.

Row notes (2) and this build's claims

LangGraph: when a node fails mid-super-step, pending checkpoint writes from nodes that completed are stored. Source · checked 2026-10-02

LangGraph: a SIGTERM drain saves a resumable checkpoint at a superstep boundary, and a drained run resumes with the same thread id. Source · checked 2026-10-02

04Deploying new code while runs are paused.
Partial

ADK's resume page says a stopped workflow must not be modified before resuming; adding or removing agents is not supported. No version-migration guidance was found.

Source · checked 2026-10-02
Partial

ADK's resume page says a stopped workflow must not be modified before resuming; adding or removing agents is not supported. No version-migration guidance was found.

Source · checked 2026-10-02
Partial

Edge changes between kept nodes are safe. Renaming or removing a node a thread is paused at or about to enter breaks its resume.

Source · checked 2026-10-02
3 Proven

On Cloud Run, a paused run finished after the service moved to a new revision.

Row notes (4) and this build's claims

Cloud Run: revisions can be rolled back, deployed gradually and given split traffic; requests being processed continue to completion. Source · checked 2026-10-02

LangGraph: the latest graph runs on every thread, including threads resuming from a checkpoint; runs are not pinned to their original code. Source · checked 2026-10-02

Agent Runtime: updating versioned fields creates an immutable revision, and traffic can be split by percentage or sent to the latest revision. Source · checked 2026-10-02

ADK: the resume page describes agent workflows, supported from ADK Python 1.16, and does not mention graph workflows or ADK 2.0. Source · checked 2026-10-02

05Approval bound to the exact proposal and a verified identity.
Partial

ADK RequestInput pauses a workflow node for human input and passes the reply on. Approver identity and a recheck before the action are not described.

Source · checked 2026-10-02
Partial

ADK RequestInput pauses a workflow node for human input and passes the reply on. Approver identity and a recheck before the action are not described.

Source · checked 2026-10-02
Partial

Human-in-the-loop middleware pauses a tool call for approve, edit, reject or respond decisions. Approver identity and a recheck before the action are not described.

Source · checked 2026-10-02
3 Proven

An approval only counts for the exact proposal it was given for.

Row notes (3) and this build's claims

ADK: tool confirmation lists DatabaseSessionService and VertexAiSessionService as not supported, and is marked Experimental. Source · checked 2026-10-02

Agent Runtime: logs show both the user's and the agent's identity when an agent acts on a user's behalf. Source · checked 2026-10-02

Cloud Run: services are private by default and callable by principals with roles such as Cloud Run Invoker. Source · checked 2026-10-02

06The model's process holds no write credentials.
Partial

Agent identity is a per-agent, least-privilege identity with certificate-bound credentials. A separate write-credential service outside the agent is not described.

Source · checked 2026-10-02
Your team

ADK tools act with the agent's own identity, such as a service account, or the controlling user's through OAuth. No separate write service is described.

Source · checked 2026-10-02
Your team

The guardrails page describes no credential handling for tools and no separate write service.

Source · checked 2026-10-02
2 Proven

Only a separate service holds GitHub and shop write credentials; the AI-facing API does not.

Row notes (1) and this build's claims
07External writes not repeated when a response is lost.
Partial

ADK resume reinstates results of tools that returned and re-runs the failed one. Tools may run more than once; ADK says to prevent duplicate runs.

Source · checked 2026-10-02
Partial

ADK resume reinstates results of tools that returned and re-runs the failed one. Tools may run more than once; ADK says to prevent duplicate runs.

Source · checked 2026-10-02
Partial

Completed task results load from the checkpoint on resume. A task that started but did not finish may run again, so make side effects idempotent.

Source · checked 2026-10-02
1 Proven

If a write's reply is lost, the retry reconciles instead of repeating it.

This build's claims
08Spend caps per run.
Partial

RunConfig max_llm_calls caps LLM calls per run (default 500). The RunConfig reference lists no token or cost limit.

Source · checked 2026-10-02
Partial

RunConfig max_llm_calls caps LLM calls per run (default 500). The RunConfig reference lists no token or cost limit.

Source · checked 2026-10-02
Partial

Model call limit middleware caps model calls per run or per thread, ending gracefully or raising an error. Its parameters are call counts.

Source · checked 2026-10-02
3 Proven

In tests, concurrent budget reservations stayed within the run's cap.

This build's claims
09An audit trail.
Partial

Agent Platform Data Access audit logs include session creates and event appends but are disabled by default. Approver-level records were not found.

Source · checked 2026-10-02
Partial

Cloud Run writes Admin Activity, Data Access and System Event audit logs for service and job operations. No documented per-run approval record was found.

Source · checked 2026-10-02
Partial

Cloud Run writes Admin Activity, Data Access and System Event audit logs for service and job operations. No documented per-run approval record was found.

Source · checked 2026-10-02
3 Proven · 1 In build · Phase 5

Each run leaves an audit trail from start to finish.

Row notes (3) and this build's claims

ADK: SessionService handles creating, retrieving, updating (appending Events, modifying State) and deleting sessions. Source · checked 2026-10-02

LangGraph: a checkpointer saves a snapshot of graph state at each super-step. Source · checked 2026-10-02

Agent Runtime: Agent Gateway logs gateway allow and deny decisions. Source · checked 2026-10-02

10Redaction before data reaches models and storage.
Partial

Agent Gateway applies Model Armor templates to agent prompts and responses and can redact or block content that violates them. Stored-data redaction is not described.

Source · checked 2026-10-02
Partial

Model Armor plugin screens user input and model output for sensitive data and blocks flagged content, showing replacement text. Stored-data redaction is not described.

Source · checked 2026-10-02
Partial

PIIMiddleware redacts, masks, hashes or blocks email, card, IP, MAC, URL or custom-regex matches in input, output or tool results. Storage paths are not described.

Source · checked 2026-10-02
3 Proven

Four planted secret types never reached stored run data or model requests.

Row notes (2) and this build's claims

ADK: the BigQuery Agent Analytics plugin has a content_formatter parameter for custom masking or formatting per event. Source · checked 2026-10-02

ADK: the BigQuery Agent Analytics plugin redacts state keys prefixed with temp: in logged state. Source · checked 2026-10-02

11Several model providers, with fallback.
Framework

ADK supports Gemini, Claude, Agent Platform-hosted models, and LiteLLM, Ollama and vLLM connectors. Model routing includes automatic failover on error.

Source · checked 2026-10-02
Framework

ADK supports Gemini, Claude, Agent Platform-hosted models, and LiteLLM, Ollama and vLLM connectors. Model routing includes automatic failover on error.

Source · checked 2026-10-02
Framework

Model fallback middleware tries alternative models when the primary model fails, with provider redundancy across OpenAI, Anthropic and others.

Source · checked 2026-10-02
2 Proven

Gemini, OpenAI and an open model each passed a role test.

This build's claims
12Tracing and analytics.
Platform

Setting GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY enables agent traces, logs and metrics. Traces open in Cloud Trace and the console.

Source · checked 2026-10-02
Framework

ADK emits OpenTelemetry spans for agent invocation, workflow execution, tool execution and model calls, exportable to Cloud Trace or an OTLP endpoint.

Source · checked 2026-10-02
Framework

LangChain and LangGraph apps have a built-in integration that sends OpenTelemetry traces to LangSmith when LANGSMITH_OTEL_ENABLED is set.

Source · checked 2026-10-02
2 Proven · 1 In build · Phase 5

An analytics row's trace id opens the matching trace in Cloud Trace.

Row notes (1) and this build's claims

ADK: the BigQuery Agent Analytics plugin trace_id is an OpenTelemetry trace ID, so rows join to Cloud Trace traces. Source · checked 2026-10-02

13What you operate and patch.
Platform

Agent Runtime is a fully-managed runtime that abstracts away the underlying infrastructure, so you focus on agent logic instead of operations.

Source · checked 2026-10-02
Partial

Google maintains and patches buildpack base images weekly; the page does not mention application dependencies. Instances are stateless, so state goes in external storage.

Source · checked 2026-10-02
Partial

Google maintains and patches buildpack base images weekly; the page does not mention application dependencies. Instances are stateless, so state goes in external storage.

Source · checked 2026-10-02
2 Proven · 1 In build · Phase 6

The whole staging stack deploys from code and can be stopped to save cost.

This build's claims
14Where state lives (data residency).
Platform

Data residency at rest is supported across Agent Runtime, evaluation, sessions and Memory Bank, alongside VPC Service Controls and CMEK.

Source · checked 2026-10-02
Your team

DatabaseSessionService stores sessions in a relational database such as PostgreSQL, MySQL, MariaDB or SQLite that you manage yourself.

Source · checked 2026-10-02
Your team

PostgresSaver stores checkpoints in a Postgres database reached through a connection string that you supply.

Source · checked 2026-10-02
2 Proven

The agent framework's session storage works on AlloyDB Omni.

Row notes (1) and this build's claims

Cloud Run: each resource resides in a region, and customer data associated with it is stored in the selected region. Source · checked 2026-10-02

15Infrastructure cost model.
Platform

Agent Runtime bills vCPU-hours and GiB-hours of allocated compute and memory, rounded to the nearest second. Idle time between prompts is not billed.

Source · checked 2026-10-02
Platform

Request-based billing charges for request processing, start and shutdown. Instance-based billing charges for the entire instance lifecycle.

Source · checked 2026-10-02
Platform

Request-based billing charges for request processing, start and shutdown. Instance-based billing charges for the entire instance lifecycle.

Source · checked 2026-10-02
2 Proven · 1 In build · Phase 6

The whole staging stack deploys from code and can be stopped to save cost.

This build's claims

"No documented support found" means the pages we read say nothing, not that the feature is impossible.

IN THE VENDORS' WORDS

Where others are stronger.

Things a baseline does that this build does not, in the vendors' own words.

Agent Runtime is fully managed and abstracts away the underlying infrastructure, so you focus on agent logic instead of operations.

Source · checked 2026-10-02

VPC Service Controls, CMEK and data residency at rest are supported across Agent Runtime, evaluation, sessions and Memory Bank.

Source · checked 2026-10-02

LangGraph checkpoints support time travel: replay from a prior checkpoint, or fork it with update_state and continue from the branch.

Source · checked 2026-10-02

After an ADK agent is deployed to Agent Platform, session management is handled automatically by the managed service.

Source · checked 2026-10-02

Agent Runtime long-running query jobs run asynchronously for up to seven days, and their status and results can be fetched later.

Source · checked 2026-10-02

Agent Gateway blocks agent connections by default unless an explicit IAM policy grants access, and can attach Model Armor screening.

Source · checked 2026-10-02

LIMITS

What this build doesn't do yet.

The limits of today's build, all of them. The phases are on the milestone line.

Next: Recorded evidence →