BoundBench

Agents (aiwaves)

Python framework for multi-agent language pipelines (SOP graphs of nodes and agents) with symbolic-learning optimizers that rewrite prompts and pipelines.

github.com/aiwaves-cn/agents · 2026-10-04 · e8c4e3c

Defense-in-depth score

0.4 / 10

Minimal

This framework executes model tool calls immediately with no approval, validation, sandbox, or run limit. Its flagship code-interpreter tool runs Open Interpreter on the host as your user, with your OpenAI key, so a hijacked or runaway agent has your whole machine. Importing it also loads a .env file and passes two environment values to Python eval, and persisted artifacts do not protect the API key. Treat it as research code, not something to run on untrusted input.

Key gaps (5)

  1. The registered code_interpreter tool runs Open Interpreter on the host as the OS user with the API key, so a hijacked agent holds the user's full local authority. C1 · Identity & least privilege
  2. Model-written code (code_interpreter, sympify, eval of env vars) executes on the host with the user's files, network, and API key; the only guard is documented as not a sandbox. C4 · Code-execution isolation
  3. A hijacked agent with the code interpreter can leak secrets and take irreversible host actions with no human in the loop (C5-WORSTCASE). C5 · Untrusted input blast radius
  4. load_dotenv() at import plus eval() of environment values lets a planted .env redirect the model endpoint or run arbitrary Python, with no trust decision (C6-REPOCONFIG). C6 · Memory, context & configuration integrity
  5. Packages installed at the model's request through the code interpreter run as the user with the full environment and API key. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

The framework has no agent identity or authorization layer. Every agent and tool runs with the developer's OS user and environment; the code-interpreter tool hands Open Interpreter the OPENAI_API_KEY and runs it in-process with full access to the host. Nothing narrows what a hijacked agent can reach.

C2 Approval gates

Minimal 0.00 / 1.00

There is no approval step anywhere. The agent loop parses the model's tool call and executes it immediately, including the code interpreter that writes and runs arbitrary code. The framework offers no hook for a human to review or reject a call, so every consequential action is unattended.

C3 Tool & action scoping

Minimal 0.05 / 1.00

Tools receive the model's arguments with no validation in code: the JSON arguments are splatted straight into the tool function. The built-in tools are general-purpose (free-form code requests, sympy expressions that sympify evaluates). The one real narrowing is that toolkits are built only from named entries in a fixed registry and the agent rejects tool names it was not given.

C4 Code-execution isolation

Minimal 0.00 / 1.00

Model-influenced code runs without isolation on several paths: the code-interpreter tool runs Open Interpreter on the host, the math tools pass model strings to sympy's eval-based parser, and two settings read from environment variables are passed to Python eval. The HumanEval scorer runs generated code in a subprocess with a function-disabling guard whose own docstring says it is not a security sandbox. Nothing contains an escape, and the host process holds the API key.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Nothing separates untrusted content from instructions. Peer agents' messages are folded into a user-role observation, tool results become the agent's next content, and retrieved knowledge-base or memory text goes straight into prompts. With the code interpreter enabled, a hijacked agent can both exfiltrate data and take irreversible actions with no human involved.

C6 Memory, context & configuration integrity

Minimal 0.05 / 1.00

Agent and shared environment long-term memory are appended to JSONL files under a relative memory/ directory that is never cleared, so content from one run persists into the next and is retrieved into routing prompts when a knowledge base is configured. The LLM module calls load_dotenv() at import, so a .env file found by python-dotenv's search can redirect the model endpoint and key, and two settings from the environment are passed to Python eval, turning a planted .env into code execution. There is no validation, provenance, or isolation on any of it.

C7 Third-party extensions

Minimal 0.05 / 1.00

The framework itself loads no plugins, MCP servers, or model files. The exposure is through the code interpreter: Open Interpreter can install whatever packages the model asks for, unpinned and unverified, and they run as the user with the full environment. That tool is off unless configured, but it is the registry's flagship tool and the chatbot example enables it.

C8 Secrets & sensitive-data protection

Minimal 0.05 / 1.00

The OpenAI key is read from the environment, but persisted artifacts do not protect the key. Logs store full prompt messages in plaintext with no redaction, the flagship example turns on litellm verbose logging, and the Trainer starts Weights & Biases by default. The key is long-lived and visible to every code-interpreter process.

C9 Audit & traceability

Minimal 0.20 / 1.00

The only record is a JSON dump of the prompt messages and the final content, written to a relative logs/ directory when a tool is called (and on other calls only if SAVE_LOGS is set). It does not record which tool ran or with what arguments, and the logger deletes all but the newest 20 files in the directory. Training mode keeps fuller trajectories, but normal runs leave no reliable audit trail.

C10 Limits & kill switch

Minimal 0.00 / 1.00

A run has no limit on steps, time, or cost. Solution.run loops until the SOP reaches its end node; the per-node max_chat_nums check is overwritten by the LLM or order transition on the next line, single-successor nodes loop forever, and the LLM call retries forever on errors. Only the training loop (max_step) and the HumanEval scorer (3-second timeout with process kill) are bounded.