C1 Identity & least privilege
Minimal 0.00 / 1.00
The web server's default network exposure and authentication are not locked down. The agent's tools run as the server's operating-system user with the server's full environment, including the LLM API key, and connect to databases with whatever stored credentials the datasource holds. There is no per-tool or per-request narrowing of authority, and no authorization layer between the agent and its tools.
C2 Approval gates
Minimal 0.20 / 1.00
The only human approval in the agent path is a confirmation prompt for connector (MCP) tools that a catalog marks as writes. Shell commands, Python, SQL, and skill scripts never require approval. The connector approval gate is not a strict boundary.
C3 Tool & action scoping
Minimal 0.15 / 1.00
The default tool set hands the model a raw shell, raw Python, and a SQL tool, all running on the server host. The SQL tool's 'SELECT only' rule is not a strict boundary and the connector beneath it executes writes and DDL. A few tools validate properly (read_file resolves real paths against an allowlist), but the skill script runner's path handling is not a complete boundary, and the shell tool makes every other check moot.
C4 Code-execution isolation
Minimal 0.42 / 1.00
By default, model-written Python and shell commands run as ordinary subprocesses on the server host, with the server's full environment (including the LLM API key) and unrestricted network access; Python code checks are explicitly switched off and the bash check is a short denylist. Docker, Podman and nerdctl backends exist and fail closed if selected, but they are off unless SANDBOX_RUNTIME is set, use a stock root container with networking on, and the skill script runner always executes on the host regardless of the setting.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
The agent reads untrusted content (uploaded files, knowledge-base documents, database rows, MCP connector results, and anything it fetches with curl) into the same context that chooses its tools, with nothing to separate data from instructions. In the same session it holds database credentials and the LLM key, can run shell commands, and can send data anywhere on the network, all without human approval.
C6 Memory, context & configuration integrity
Minimal 0.05 / 1.00
Skills are folders of instructions and scripts on disk; when one is loaded its SKILL.md is injected into the system prompt with an order to follow it strictly, and the agent's unsandboxed shell can write new skill folders into that same directory, which serves every user. Past conversation turns are reloaded into later rounds, and knowledge-base documents feed retrieval. Conversations are keyed by a user name, and skills and knowledge spaces are shared.
C7 Third-party extensions
Minimal 0.00 / 1.00
The shell tool's own description invites the model to run pip, apt, and curl, so the agent can install and execute third-party packages on the host whenever it decides to, without asking. Skills imported from GitHub are fetched from a branch tip with no pinning or integrity check, and their scripts run as host subprocesses with the server's full environment. MCP connectors are remote servers the user adds explicitly; custom connectors bypass the catalog entirely.
C8 Secrets & sensitive-data protection
Minimal 0.20 / 1.00
The setup wizard writes its config, which can hold the LLM API key, to a file created with owner-only permissions, and secret-named arguments are masked in approval prompts. Elsewhere, database passwords are stored as plain text in the metadata database, connector-credential encryption is not robust, and every code and skill subprocess receives the server's full environment, so the model can read any key it wants. A local tracer records LLM replies and tool results by default; there is no third-party telemetry.
C9 Audit & traceability
Minimal 0.40 / 1.00
A tracer that is on by default writes a span for every agent action, including the model's reply and the tool output, to a local JSONL file and an SQLite store, and the conversation history with every step is saved at the end of each turn. These records do not reliably say who asked, approvals are not recorded, spans are batched and flushed late, and everything sits on the same host the agent's unsandboxed shell can edit.
C10 Limits & kill switch
Minimal 0.45 / 1.00
The main agent is capped at 30 reasoning steps, shell commands at 30 seconds, Python at 60 seconds, and each sub-agent at 15 steps and 10 minutes, with at most three sub-agents per dispatch and no nested dispatch. There is no token or cost cap and no overall time limit, and each dispatch hands sub-agents fresh budgets. Closing the stream cancels the agent and kills the local process tree, but the skill script runner leaves its subprocess running when it times out.