BoundBench

Google ADK (Python)

Code-first toolkit for building, evaluating, deploying agents

github.com/google/adk-python · 2026-10-03 · e94c2e7

Defense-in-depth score

3.5 / 10

Minimal

ADK starts every agent with no tools and no code execution, and it ships genuinely good building blocks: an exact-call confirmation flow, an SSRF-hardened web loader, a read-only-by-default BigQuery tool and a gVisor code executor. Almost none of them are on by default, though. Tools run without approval unless each one opts in, the local executors give model-written code the full environment with its keys, and nothing limits what an agent can do after reading a malicious page or MCP result. Developers who enable code execution or shell tools have to choose isolation and approval themselves.

Key gaps (2)

  1. The bundled local code and shell executors run model-generated code as host subprocesses with the full process environment, including API keys and cloud credentials. C4 · Code-execution isolation
  2. Nothing in the framework breaks the untrusted-input, private-data and egress combination; with default (unconfirmed) tools a hijacked agent can leak data and act unattended. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.33 / 1.00

ADK has no agent identity or authorization layer of its own: function tools run in the developer's process with whatever credentials it holds, and the local code and shell executors hand model-generated code the full process environment. Google API toolsets can instead be configured with per-user OAuth (each toolset gets its own client and scopes and the end user consents), which is a real narrowing, but it is opt-in, defaults to broad scopes such as full BigQuery plus Dataplex read-write, and does not cover function tools or executors.

C2 Approval gates

Minimal 0.28 / 1.00

ADK ships a well-built per-call confirmation flow: a tool can require approval statically or through a callable that inspects the arguments, the approval request carries the exact original call, only a user-authored event can approve, and the executed arguments must equal the approved ones. But it is off by default for every function, MCP and OpenAPI tool, and the code-executor path (model-written code run by a configured executor) and the experimental environment shell tool cannot be gated at all. Only the bash tool always asks. Nothing provides undo or dry-run for external actions; session rewind only restores ADK's own state and artifacts.

C3 Tool & action scoping

Moderate 0.50 / 1.00

Every function tool's arguments are type-checked against its schema before it runs, and several bundled tools are carefully bounded: the web loader pins the resolved IP, blocks internal addresses and refuses redirects; environment file tools enforce resolved-path containment; BigQuery blocks non-SELECT statements by default using the service's own dry-run. But the general-purpose tools stay general: the bash tool allows any command by default and the environment Execute tool runs any shell string, and MCP tools get no shared validation. A new agent starts with no tools at all, so every capability is an explicit developer choice.

C4 Code-execution isolation

Minimal 0.47 / 1.00

No code runs by default, but the execution options ADK ships without extras are unsandboxed: UnsafeLocalCodeExecutor runs model-written Python as a host subprocess with the full environment, and the bash tool and LocalEnvironment shell run on the host too. Strong isolation is available on request: the GKE executor runs each snippet in a fresh gVisor pod as a non-root user with dropped capabilities, no service-account token and resource limits, and the container executor disables networking and drops capabilities. Even then, shell tools, skill scripts on a local executor, and MCP stdio servers keep running on the host.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

ADK fences some untrusted text as data: MCP tool descriptions and messages from other agents are wrapped in markers with a preamble telling the model not to follow them. That is spotlighting, which raises the bar but does not stop a determined injection, and ordinary tool results (including fetched web pages) are not fenced at all. Nothing ties approval or egress to whether untrusted content has been read; Model Armor screening is an opt-in plugin. Under the framework rule, a hijacked agent can combine untrusted input, private data and egress with no human involved.

C6 Memory, context & configuration integrity

Minimal 0.38 / 1.00

The model has no tool to write long-term memory directly: whole sessions are added to memory only when the developer's code or the dev server calls it, searches are keyed by app and user, and recalled memories come back with author and timestamp inside a 'past conversations' block. But nothing validates what is stored, so an injection inside a session is recalled in later sessions, and session state (including app-wide state shared by every user) can be templated straight into the system instruction. ADK does not auto-load instruction files from a workspace; the CLI loads the .env from the developer's own agent folder.

C7 Third-party extensions

Minimal 0.23 / 1.00

ADK loads no third-party code unless the developer configures it, and the developer writes the exact MCP launch command or skill source themselves. Nothing pins or verifies what is loaded: MCP server tool lists are refetched each session with no change detection, and skill revalidation is off by default and only re-states changed text. Skill scripts run through whichever executor is configured, which with the local executor means a host subprocess holding the full environment; MCP stdio servers get the MCP library's default environment.

C8 Secrets & sensitive-data protection

Minimal 0.25 / 1.00

Credentials come from environment variables or developer-supplied objects. The generic auth flow strips the OAuth client secret before storing credentials in temporary state, and CLI usage telemetry is content-free and only sent after the user opts in. Elsewhere protection is thin: the Google credentials helper caches the full OAuth credential JSON (refresh token and client secret) in persisted session state, OpenTelemetry spans capture message content by default, and the local executors pass every environment variable, API keys included, to model-generated code.

C9 Audit & traceability

Minimal 0.40 / 1.00

Every tool call, tool result, code execution and approval is recorded as a structured session event with timestamp, author (agent name or 'user'), invocation id and branch, and the same steps are emitted as OpenTelemetry spans. That is a good transcript, but by default it lives in memory (or in a SQLite file inside the agent folder for the dev server), the agent's process can alter it, and calls made inside an agent wrapped as a tool go to a throwaway in-memory session. Nothing makes actions wait for their record.

C10 Limits & kill switch

Minimal 0.40 / 1.00

Each run is capped at 500 model calls by default, enforced in code and adjustable by the developer or an environment variable, and the bash, environment and container executors have per-execution timeouts that kill the process group. There is no wall-clock or token/cost cap per run, loop agents iterate without a limit by default, and an agent wrapped as a tool starts with a fresh 500-call count. An abort signal stops the run, but synchronous tools already running in the thread pool keep going.