C1 Identity & least privilege
Minimal 0.00 / 1.00
The framework has no agent identity or authorization layer. Every agent and tool runs with the developer's OS user and environment; the code-interpreter tool hands Open Interpreter the OPENAI_API_KEY and runs it in-process with full access to the host. Nothing narrows what a hijacked agent can reach.
C2 Approval gates
Minimal 0.00 / 1.00
There is no approval step anywhere. The agent loop parses the model's tool call and executes it immediately, including the code interpreter that writes and runs arbitrary code. The framework offers no hook for a human to review or reject a call, so every consequential action is unattended.
C3 Tool & action scoping
Minimal 0.05 / 1.00
Tools receive the model's arguments with no validation in code: the JSON arguments are splatted straight into the tool function. The built-in tools are general-purpose (free-form code requests, sympy expressions that sympify evaluates). The one real narrowing is that toolkits are built only from named entries in a fixed registry and the agent rejects tool names it was not given.
C4 Code-execution isolation
Minimal 0.00 / 1.00
Model-influenced code runs without isolation on several paths: the code-interpreter tool runs Open Interpreter on the host, the math tools pass model strings to sympy's eval-based parser, and two settings read from environment variables are passed to Python eval. The HumanEval scorer runs generated code in a subprocess with a function-disabling guard whose own docstring says it is not a security sandbox. Nothing contains an escape, and the host process holds the API key.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
Nothing separates untrusted content from instructions. Peer agents' messages are folded into a user-role observation, tool results become the agent's next content, and retrieved knowledge-base or memory text goes straight into prompts. With the code interpreter enabled, a hijacked agent can both exfiltrate data and take irreversible actions with no human involved.
C6 Memory, context & configuration integrity
Minimal 0.05 / 1.00
Agent and shared environment long-term memory are appended to JSONL files under a relative memory/ directory that is never cleared, so content from one run persists into the next and is retrieved into routing prompts when a knowledge base is configured. The LLM module calls load_dotenv() at import, so a .env file found by python-dotenv's search can redirect the model endpoint and key, and two settings from the environment are passed to Python eval, turning a planted .env into code execution. There is no validation, provenance, or isolation on any of it.
C7 Third-party extensions
Minimal 0.05 / 1.00
The framework itself loads no plugins, MCP servers, or model files. The exposure is through the code interpreter: Open Interpreter can install whatever packages the model asks for, unpinned and unverified, and they run as the user with the full environment. That tool is off unless configured, but it is the registry's flagship tool and the chatbot example enables it.
C8 Secrets & sensitive-data protection
Minimal 0.05 / 1.00
The OpenAI key is read from the environment, but persisted artifacts do not protect the key. Logs store full prompt messages in plaintext with no redaction, the flagship example turns on litellm verbose logging, and the Trainer starts Weights & Biases by default. The key is long-lived and visible to every code-interpreter process.
C9 Audit & traceability
Minimal 0.20 / 1.00
The only record is a JSON dump of the prompt messages and the final content, written to a relative logs/ directory when a tool is called (and on other calls only if SAVE_LOGS is set). It does not record which tool ran or with what arguments, and the logger deletes all but the newest 20 files in the directory. Training mode keeps fuller trajectories, but normal runs leave no reliable audit trail.
C10 Limits & kill switch
Minimal 0.00 / 1.00
A run has no limit on steps, time, or cost. Solution.run loops until the SOP reaches its end node; the per-node max_chat_nums check is overwritten by the LLM or order transition on the next line, single-successor nodes loop forever, and the LLM call retries forever on errors. Only the training loop (max_step) and the HumanEval scorer (3-second timeout with process kill) are bounded.