BoundBench

goose

Extensible open-source AI agent (CLI + desktop) that installs, executes, edits and tests

github.com/aaif-goose/goose · 2026-10-03 · 591edd4

Defense-in-depth score

2.1 / 10

Minimal

goose ships in Auto mode, which runs every shell command, file write and extension call without asking, directly on the host with the user's full environment and credentials. A per-call Approve mode exists but is opt-in, and its enforcement does not cover every path. The most serious issue is that plugins committed to a repository's .agents/plugins folder are auto-enabled, so opening goose in a cloned repo can launch that repo's MCP servers and hook scripts with no prompt. Run it in a disposable VM or container and switch to Approve mode before pointing it at untrusted code.

Key gaps (5)

  1. Plugins under the working directory's .agents/plugins are auto-enabled, and their hook commands (including SessionStart) and MCP servers run on the host with no trust prompt. C6 · Memory, context & configuration integrity
  2. A cloned repo's plugin MCP servers and hooks are launched by default without user consent. C7 · Third-party extensions
  3. In the default Auto mode a hijacked session can exfiltrate secrets and take irreversible actions with no human in the loop. C5 · Untrusted input blast radius
  4. Model-written shell commands run unsandboxed on the host with the full inherited environment. C4 · Code-execution isolation
  5. Every tool runs with the user's ambient credentials; nothing narrows or checks authority. C1 · Identity & least privilege

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

goose runs as the logged-in user with no narrowing of that authority. The shell tool starts bash with the full inherited environment, so every API key, cloud credential chain, SSH agent socket and gh/git login available to the user is available to model-written commands, and no subprocess anywhere clears its environment. There is no per-tool identity or authorization layer; in the default Auto mode every tool call is allowed. A hijacked session therefore holds the user's entire account across every service they are logged into.

C2 Approval gates

Minimal 0.38 / 1.00

Approval is off by default: goose's default mode is Auto, which approves every tool call (shell, file writes, extension tools) without asking, and the first-run setup preselects Auto. An Approve mode exists that shows the exact tool name and arguments for each call and offers Allow, Always Allow, Deny and Cancel, with user-set per-tool permission levels. Even in Approve mode, the gate does not cover every path, plugin hooks run outside it, and Always Allow covers a whole tool such as the shell. Nothing in goose offers checkpoints or undo for shell side effects.

C3 Tool & action scoping

Minimal 0.13 / 1.00

The default tool set is as broad as it gets: an arbitrary shell command string, file write/edit that accepts any absolute path on the machine, and an image reader that fetches any http(s) URL. There is no path containment, URL allowlist, or command allowlist on any built-in tool. Extensions can be toggled off as a whole, but the Developer extension with shell and write tools is on by default.

C4 Code-execution isolation

Minimal 0.20 / 1.00

Model-written shell commands run directly on the host as the user through bash -c, with no container, OS sandbox or filtering. A --container flag can run stdio and built-in MCP extensions through docker exec in a container the user provides, but it does not cover every execution path. Repo-supplied plugin hooks and MCP servers also execute on the host.

C5 Untrusted input blast radius

Minimal 0.15 / 1.00

goose reads untrusted content routinely (repository files, shell and web output, MCP tool results, and MCP server instructions that are placed in the system prompt) and nothing structural limits what a hijacked session can do. In the default Auto mode the agent can read secrets, send them anywhere through the shell or the URL-fetching image tool, and take irreversible actions with no human involved. An optional prompt-injection pattern scanner exists but is off by default and only detects; the egress inspector only logs.

C6 Memory, context & configuration integrity

Minimal 0.05 / 1.00

A cloned repository can change goose's behavior with no prompt. Plugins placed under .agents/plugins in the working directory are discovered and auto-enabled, their MCP servers are launched, and their hooks.json commands (including SessionStart) run through sh on the host. .goosehints and AGENTS.md from the project are loaded silently into the system prompt, and the model can write those files (or a new plugin) with its unrestricted write and shell tools, so an injection can persist into later sessions and to anyone else who clones the repo.

C7 Third-party extensions

Minimal 0.07 / 1.00

Third-party code runs with the agent's full access and little verification. User-configured MCP extensions are typically npx/uvx packages launched unpinned; goose checks them against the OSV malicious-package list, but that check only covers npx/uvx and fails open otherwise. Worse, plugins committed to a project's .agents/plugins directory are enabled on discovery, so opening goose in a cloned repo launches that repo's MCP servers and hook scripts without consent. Extension processes run as the same user with the inherited environment.

C8 Secrets & sensitive-data protection

Minimal 0.20 / 1.00

Provider API keys go to the OS keychain by default, falling back to an owner-only (0600) file. Beyond that there is little protection: shell commands inherit every environment secret and their output goes to the model, nothing redacts secrets from tool results, logs or saved transcripts (the only redaction is for telemetry error strings), and the CLI writes full LLM request payloads to local log files by default. Telemetry is opt-in.

C9 Audit & traceability

Moderate 0.50 / 1.00

Every message, including tool calls with arguments and results, is written per message to a SQLite session database in the user's data directory, and sub-agent sessions are linked to their parent. That gives a usable local transcript. It is not tamper-evident: the agent's own shell runs as the same user and can edit or delete it, and there is no actor attribution separating approver from agent or standard audit export by default.

C10 Limits & kill switch

Minimal 0.40 / 1.00

The main loop stops after 1,000 turns and shell commands get a 300-second default timeout, but the model sets timeout_secs itself and a value of 0 disables the timeout. Sub-agents get a fresh 25-turn budget each (model-adjustable), cannot delegate further, and at most 5 run in the background. There is no cost or token cap. Ctrl-C cancels the loop and kills the shell process, but not its process group, so backgrounded children can survive.