C1 Identity & least privilege
Minimal 0.38 / 1.00
Codewhale runs as you and does not get a scoped identity of its own, but it narrows what its child processes inherit: every shell command, test runner, language server and MCP server starts from a cleared environment with a short allowlist of non-secret variables, so API keys and tokens in your shell are not passed on. Your home directory is still reachable, though, so credential files on disk (cloud configs, the gh login, SSH keys) remain usable by any command you approve, and on Linux nothing confines those commands. Widening (Full Access, trust mode) is an explicit user choice.
C2 Approval gates
Moderate 0.63 / 1.00
In the default Ask posture every shell command needs your approval, except a short, carefully parsed allowlist of read-only commands (cat, ls, rg, git status/log/diff and read-only gh queries) that refuses pipes, redirects, substitutions and unknown options. Web fetch and search also ask every time, and MCP tools with side effects ask. The approval card shows the exact command or file content. The main exception is file edits: inside a git repository, writes to project files run without a prompt (credential files, .git and .codewhale are excluded), relying on git and snapshots for undo. Full Access is one Shift+Tab away and turns prompts off; a repository cannot loosen any of this.
C3 Tool & action scoping
Minimal 0.40 / 1.00
The file tools check that paths stay inside the workspace after resolving symlinks, and web fetches go through an SSRF guard that blocks private and cloud-metadata addresses, pins DNS and rechecks redirects. But the main tool is a general shell that accepts any command string; only the auto-approval classifier is an allowlist. Shell, file write and network tools are all available by default in Work mode (Plan mode is read-only), and on Linux a misused approved command reaches the whole machine.
C4 Code-execution isolation
Moderate 0.50 / 1.00
On Linux, the default scored here, shell commands run as ordinary child processes with no OS sandbox: bubblewrap is used only if you set prefer_bwrap = true, and if bubblewrap is missing it silently runs unsandboxed. When enabled, bubblewrap gives a read-only view of the disk, write access only to the workspace and temp dirs, and no network, and test runners use the same confinement, but MCP servers and language servers still run on the host. On macOS a Seatbelt profile is applied automatically, which would score higher. Without the sandbox, a command reaches your whole home directory.
C5 Untrusted input blast radius
Moderate 0.50 / 1.00
Codewhale does not try to detect prompt injection; its protection is that the things a hijacked agent would need are behind approval. Web fetch, web search and any non-read-only shell command ask you each time, regardless of what the agent has read, so sending data out or running destructive commands needs a person. What a hijacked agent can do unattended is read files (including, on Linux, anything the read-only shell allowlist can reach) and edit project files in a git repository. Tool output is masked for credential-shaped values before reaching the model.
C6 Memory, context & configuration integrity
Minimal 0.25 / 1.00
Codewhale is careful with repository-controlled settings: a project's .codewhale/config.toml can only tighten approval, sandbox and shell settings and cannot change endpoints or MCP config, project MCP servers and skills load only after you trust the folder in your own config, and project hooks additionally need a content-hash approval. Long-term memory is off by default and, when on, the model can only propose candidates for you to review. However, AGENTS.md and similar instruction files load silently, and the agent can rewrite them without a prompt in a git repo, so injected instructions can persist into later sessions. A workspace .env is also read at startup, before any trust decision, and can supply model-provider and sandbox-service API keys when you have not configured your own.
C7 Third-party extensions
Minimal 0.38 / 1.00
No third-party code runs by default; the bundled Computer Use plugin is off until reviewed. Plugins get a strong review: each is bound to a hash of its content and declared capabilities, any change forces re-review, and the plugin runtime host runs sandboxed. MCP servers are weaker: servers from your config run whatever command you set, with no pinning or integrity check, and the model can ask to install a server from the public MCP Registry (pinned to a version) or start an arbitrary MCP server command, each behind a per-call approval. MCP servers run as unsandboxed host processes with a scrubbed environment. A trusted project's .codewhale/mcp.json can add servers after the generic folder-trust prompt.
C8 Secrets & sensitive-data protection
Minimal 0.47 / 1.00
Credential-shaped values and configured keys are masked in tool output before it reaches the model and before it is written to saved sessions, and turning that masking off needs a confirmed, receipt-bound opt-out. Child processes never receive secret-shaped environment variables. API keys are stored in a file with owner-only permissions by default (the OS keyring is opt-in). Usage telemetry is on by default and documented as content-free.
C9 Audit & traceability
Moderate 0.57 / 1.00
Each session is saved under ~/.codewhale/sessions with every tool call and result, and a per-session approval receipt log records each approval request and its outcome, including whether a person, a session rule or the posture decided it. These records sit outside the workspace but are ordinary files the agent's own (unsandboxed, on Linux) commands could alter. A separate structured tool-audit log exists only when an environment variable is set. There is no tamper-evidence or off-host export.
C10 Limits & kill switch
Minimal 0.40 / 1.00
By default a turn has no step limit and no wall-clock limit, there is no spend ceiling, and a goal in Operate mode keeps continuing until it is done or you stop it. Individual steps are bounded: shell commands time out after 120 seconds by default, model streams after 30 minutes, and each sub-agent run after 30 minutes, with nesting depth 3 and up to 64 concurrent sub-agents. Pressing Esc cancels the turn and shell commands are killed as a process group. You can configure step and time limits.