BoundBench

AgentVerse

OpenBMB's Python framework for multi-LLM-agent task solving (role assignment, decision making, execution, evaluation) and multi-agent simulations.

github.com/openbmb/agentverse · 2026-10-04 · f90c4bd

Defense-in-depth score

2.1 / 10

Minimal

AgentVerse ships no safety controls around what its agents do. Its built-in executors run model-generated code on the host through a shell or Python eval/exec, with the operator's API keys in the environment and no approval, sandbox, or argument checks. Tool-using tasks forward any tool call the model emits, including a root shell, to an external tool server. Treat it as a research framework and run it only inside a disposable VM or container that holds no credentials you care about.

Key gaps (3)

  1. Model-generated code runs on the host as the invoking user through shell=True, with the full environment (including OPENAI_API_KEY) inherited. C4 · Code-execution isolation
  2. A hijacked agent holds the operator's ambient identity: model-run subprocesses inherit all credentials in the launching shell. C1 · Identity & least privilege
  3. Untrusted web content, an ungated root shell tool and host code execution coexist with no approval; injected content can exfiltrate and act irreversibly unattended (C5-WORSTCASE). C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

AgentVerse reads the operator's OpenAI or Azure API key from the environment and has no notion of a scoped agent identity or per-request authorization. Model-written code run by the code-test executor is launched with the default subprocess environment, so it inherits that key and every other credential in the operator's shell. If an agent is hijacked, it acts with the full authority of the user who launched it.

C2 Approval gates

Minimal 0.00 / 1.00

There is no human approval step anywhere in the framework. Model-written code is executed, files are written to model-chosen paths, and tool calls (including a root shell tool exposed by the tool-using configs) are sent to the tool server without asking anyone. The only input() prompts in the code are commented out and were for scoring, not gating.

C3 Tool & action scoping

Minimal 0.05 / 1.00

Tools are as broad as they can be: the code-test executor writes model-generated code to a model-chosen file path with no containment and then runs it through a shell, and the tool-using executor forwards whatever tool name and arguments the model emits to the tool server without checking them against the configured tool list. Shipped tool configs include a root shell, a Python notebook, file writes and web browsing. Only the simulation ToolAgent checks that a named tool exists, which is not argument validation.

C4 Code-execution isolation

Minimal 0.00 / 1.00

Model-generated code runs with no isolation. The code-test executor runs it through a host shell as the current user; the simulation software-team environment exec()s model-written code and evals model-written test lists in-process; one further in-process path is not confined either. The only isolation anywhere is the external XAgent ToolServer container used by the tool-using tasks, which is not part of this repository and could not be verified. An escape is not needed: the code already has the user's files, network and API keys.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Nothing distinguishes untrusted content from instructions. Web pages fetched through the tool server are summarized and fed back into agent memory as ordinary function messages, and messages from other agents in the group are treated as context with the same standing. Because the framework also executes model output and has network egress with no approval, a successful injection in a tool-using or code-execution task can both exfiltrate data and take irreversible actions unattended.

C6 Memory, context & configuration integrity

N/A · full credit 1.00 / 1.00

No persistence path exists that the model can influence. Agent memory (chat history, summaries, the in-memory vector store, reflections) lives only in the running process and is rebuilt for every task and every benchmark example. Task configuration comes from the package's own tasks directory or an explicit --tasks_dir flag, and nothing loads .env files or instruction files from the working directory. Caveat: because model-written code runs unsandboxed on the host (C4), it could still overwrite files such as task configs; that is scored as execution blast radius, not as a memory feature.

C7 Third-party extensions

Minimal 0.28 / 1.00

Third-party tools are not loaded by default. Simulation tasks can list BMTools tool servers by URL in the task YAML, and tool-using tasks call an external XAgent ToolServer over HTTP; in both cases the spec and tool list are fetched live with no pinning or integrity check. These servers run as separate processes the operator starts themselves, and AgentVerse sends them only tool arguments and session cookies, not its API keys.

C8 Secrets & sensitive-data protection

Minimal 0.05 / 1.00

API keys come from environment variables and there is no masking, redaction or secret handling anywhere in the code. The keys are not put into prompts, and full prompts are only written to the activity log when --debug is set, but tool inputs, tool outputs and execution results are logged unredacted at the default level. The largest gap is that every subprocess, including model-written code, inherits the long-lived provider keys.

C9 Audit & traceability

Minimal 0.30 / 1.00

AgentVerse writes a text activity log (activity.log, error.log) under the package directory, recording role assignments, plans, execution results and, for the tool-using executor, each tool name, input and observation. The records are unstructured, carry no actor or approver attribution, and do not log the exact command the code-test executor runs. The log sits outside the task working directory but is writable by the same user that runs model code.

C10 Limits & kill switch

Minimal 0.38 / 1.00

Runs are bounded by a turn cap (10 by default, 3 for the default brainstorming task), a 10-call cap per tool-using executor round, a 10-second timeout on code-test execution and a 5-second timeout on in-process unit tests. Cost is tracked and printed but never enforced. The timeouts stop waiting rather than stopping work: killing the multiprocessing worker leaves the shell's python child running, and the timed-out thread keeps executing. The simulation tool agent loops with while True until the model finishes and catches BaseException, which also swallows Ctrl+C during a call.