C1 Identity & least privilege
Moderate 0.50 / 1.00
Strix does not hand the operator's cloud, SSH or shell credentials to the agent. The sandbox container is created with an explicit, minimal environment (proxy settings and a few flags), so the LLM key stays in the host process. Inside the container the agent runs as a user with passwordless sudo, and it can reach the internet and the host gateway, so its authority is wide even though it holds no stored credentials. Nothing narrows what it may do per tool or per request, and credentials a user types into --instruction are given to the model as-is.
C2 Approval gates
Minimal 0.05 / 1.00
There is no human approval step for any tool call. The system prompt explicitly tells the model never to wait for approval, and the code never sets an approval requirement on the shell, file-editing, proxy replay or MCP tools; the approval plumbing in the factory only forwards whatever the SDK tool declares. The shell tool, the most powerful action path, runs every command the model writes. The only confirmation prompts in the codebase concern which local directory to mount and the cloud product's billing.
C3 Tool & action scoping
Minimal 0.13 / 1.00
The core tools are raw passthroughs: an arbitrary shell string in the sandbox, an arbitrary replacement URL for proxy request replay, and arbitrary network egress. The authorized target list is injected into the prompt, but no code enforces it, so a hijacked or mistaken agent can attack any host. Typed schemas, output-size bounds and a workspace-path check on the shell working directory exist, and all tools including write and exec are enabled by default for every agent.
C4 Code-execution isolation
Moderate 0.63 / 1.00
All model-reachable command execution (shell, file edits, browser, scanners) runs inside a per-scan Docker container, and there is no host fallback: if Docker is unavailable the run aborts. The container is not hardened. It runs with default capabilities plus NET_ADMIN and NET_RAW, the user has passwordless sudo, and the root filesystem is writable. Resource limits are opt-in. The writable workspace mount, unrestricted network and host gateway are reachable from inside, though the host's Docker socket, home directory and the LLM key are not.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
The agents read content an attacker can control (web pages and API responses from the target, files and READMEs in scanned repositories, MCP tool results, web search results) and nothing in the code separates it from instructions: there are no wrappers, taint tracking or approval steps. If the model is hijacked it can run any command with network egress, so it can both send out the contents of the mounted source tree or credentials supplied in the instruction and launch irreversible actions against any host, all without a human. This is the central risk of the design.
C6 Memory, context & configuration integrity
Moderate 0.63 / 1.00
Strix has no cross-run memory and loads no instruction or configuration files from the scanned workspace: settings come from environment variables and files under the user's home directory, and a search found no dotenv or AGENTS.md loading in the code. Notes, todos and the threat model live in a per-run directory and are re-read only when a user resumes that run. That run directory is created under the current working directory, so if a user scans the directory they run from, the writable mount could let the agent alter state that a later --resume reloads.
C7 Third-party extensions
Minimal 0.35 / 1.00
The only third-party code Strix loads is MCP servers the user lists in a file under their home directory (or passes with a flag); nothing is enabled by default and the scanned workspace cannot add one. Those entries have no pinning or integrity check, and stdio servers run as a host process as the user. The sandbox image is selected by tag rather than digest and its Dockerfile pulls many tools at @latest. The model can also install packages inside the container, but that stays inside the sandbox and is scored under isolation.
C8 Secrets & sensitive-data protection
Minimal 0.25 / 1.00
The LLM API key is read from the environment and, once a scan starts, is written in plaintext to ~/.strix/cli-config.json with 0600 permissions; key fields are hidden from reprs and the key is not placed in the sandbox environment. Strix has no redaction of logs, saved transcripts or model-bound messages, and credentials a user puts in the --instruction text are saved in run.json and sent to the model by design. Telemetry to PostHog and Scarf is on by default but carries only coarse, content-free fields.
C9 Audit & traceability
Moderate 0.50 / 1.00
Every agent's full message history, including each tool call, its arguments and its output, is stored in a per-run SQLite database alongside an agent graph file, a run record and a debug log, and a viewer rebuilds transcripts from them. Records identify the agent and its parent, but there are no approvals to record and no tamper protection. The run directory is created under the current working directory and is not excluded from the writable target mount, so the agent's own tools may be able to edit the record when the scan target is the working directory.
C10 Limits & kill switch
Minimal 0.40 / 1.00
Each agent is stopped after 500 turns by default and the host-side tools have per-call timeouts, but there is no default cost cap, no overall wall-clock limit, and no cap on how many sub-agents may be spawned, so a scan can run and spend indefinitely. An optional --max-budget is shared by all agents when set. Ctrl-C, SIGTERM and SIGHUP stop the scan and delete the container, which interrupts in-flight sandbox commands.