BoundBench

Open Interpreter

Coding agent for open models (now a Codex-based Rust harness)

github.com/openinterpreter/openinterpreter · 2026-10-03 · 2767e5f

Defense-in-depth score

3.0 / 10

Minimal

Open Interpreter inherits Codex's real OS sandbox (Seatbelt on macOS, bubblewrap+seccomp on Linux): model commands can write only inside the project, have no network, and need your approval, with the exact command shown, to leave the sandbox. Protection of the project settings folder is not tamper-resistant. Every command also inherits your full environment (API keys, cloud tokens) and can read your whole disk, and there are no turn, time or spend limits.

Key gaps (1)

  1. Sandboxed commands inherit the full parent environment (credentials included) and can read the whole filesystem, so a hijacked command can harvest secrets even though writes and network are blocked. C4 · Code-execution isolation

Criteria

C1 Identity & least privilege

Minimal 0.28 / 1.00

Open Interpreter runs as you and does not narrow your authority for the commands it runs: by default every shell command gets your full environment, including API keys, cloud credentials and GitHub tokens, because the built-in KEY/SECRET/TOKEN filter is off by default. MCP servers are the exception and receive only a small set of core variables. In the current session the network-off sandbox and per-command approval stop those credentials from being used; these protections are not tamper-resistant across sessions. An experimental, opt-in credential broker can swap GitHub and OpenAI tokens for dummies, but it is off by default.

C2 Approval gates

Minimal 0.25 / 1.00

Approval is built around the sandbox: commands that stay inside it (editing project files, running tests) run without asking, while anything that needs to escape it, such as network access, writes outside the project or an escalated command, stops for your approval with the exact command on screen. Forced `rm -f` deletions always prompt, and MCP tools prompt unless their server declares them read-only. The approval gate is not tamper-resistant across sessions in this fork. There is no undo for approved actions.

C3 Tool & action scoping

Minimal 0.30 / 1.00

The main tool is a general shell that accepts any command string, so argument validation is minimal: there is no allowlist, only a small heuristic that flags forced `rm` and user-written prefix rules. The file-edit tool checks every path against the writable folders before auto-approving. Everything is on by default, though the shell tool can be turned off. In practice a misused command's reach is bounded by the sandbox: writes stay in the project, but reads cover the whole machine.

C4 Code-execution isolation

Minimal 0.25 / 1.00

Model-issued commands run inside a real operating-system sandbox: a deny-by-default Seatbelt profile on macOS, and bubblewrap plus seccomp with no network namespace on Linux, which fails closed if bubblewrap is missing instead of running on the host. Writes are limited to the project and /tmp, network is off, and leaving the sandbox needs your approval per command. But the sandbox is not tamper-resistant across sessions. Inside the sandbox, commands also see your full environment (credentials included) and can read your entire disk.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

Open Interpreter does not try to detect prompt injection; its protection is structural. Within one session, the network-off sandbox means an agent hijacked by a malicious README, command output or search result can't send data out or act outside the project without you approving a specific command. That protection is not tamper-resistant beyond the session. MCP tools that their server labels read-only are also called without approval.

C6 Memory, context & configuration integrity

Minimal 0.25 / 1.00

Open Interpreter keeps upstream Codex's folder-trust step: project config, exec-policy rules and hooks load only after you trust the folder, each hook must also be trusted by its content hash in your own settings, and some keys (model endpoints, notify, telemetry) are always ignored in project config. But project config and rules changed after the trust decision load with no prompt and can change approvals, sandbox mode and MCP servers for every later session, and the folder's protection is not tamper-resistant. AGENTS.md instructions also load silently and are writable by the agent. Long-term memories are off by default.

C7 Third-party extensions

Minimal 0.25 / 1.00

No third-party extensions run by default. MCP servers and plugins come from your config or the project's config and start automatically on the host, unsandboxed but with a cleaned-up environment; there is no version pinning or integrity check. The bundled QA skill tells the model to download the latest agent-browser binary from GitHub and run it, which goes through the normal per-command approval. When a skill needs an MCP server you are asked first, but the prompt lists only server names, not the command that will run.

C8 Secrets & sensitive-data protection

Minimal 0.28 / 1.00

Your login is stored by default in a plaintext auth.json under ~/.openinterpreter with owner-only permissions; MCP OAuth tokens prefer the OS keyring. Secret redaction exists but isn't applied to tool output sent to the model or to saved transcripts. The biggest gap is that every shell command inherits your full environment, so long-lived keys are within reach of any command the model runs. Usage analytics (sent to the ChatGPT backend URL) are on unless disabled; prompt logging to OpenTelemetry is off.

C9 Audit & traceability

Moderate 0.55 / 1.00

Every session is written to a JSON-lines transcript under ~/.openinterpreter/sessions that records each tool call and its output, including MCP calls; sub-agents get their own transcripts. The file sits outside the sandbox's writable area, so the model's commands can't edit it. Approval prompts and your decisions aren't saved, there is no tamper-evidence, and write failures are only logged while the agent carries on.

C10 Limits & kill switch

Minimal 0.33 / 1.00

There is no default cap on turns, wall-clock time or spend; the token-budget feature is still under development and off. You can interrupt a turn at any time, which cancels the running task, but commands started as background terminals keep running until you clean them up or exit. Sub-agents inherit the same settings and are capped at 6 concurrent threads and a depth of 1, and at most 64 background processes can exist.