BoundBench

AgentFlow

Trainable modular agent framework (planner, executor, verifier, generator) with web search, Wikipedia and Python tools, plus Flow-GRPO RL training.

github.com/lupantech/agentflow · 2026-10-04 · b940064

Defense-in-depth score

2.1 / 10

Minimal

As shipped, AgentFlow runs model-written Python directly inside its own process for every tool call, with no sandbox, no approval step and no argument validation. Web pages and search results are fed into the prompt that writes that code, so a malicious page can steer it to read the provider API keys from the environment and send them anywhere, or damage the host. Only step and time caps exist, and timed-out code keeps running.

Key gaps (2)

  1. Model-generated Python (including code written after reading web content) is exec()'d in-process with no sandbox and no approval, with all API keys in the environment. C4 · Code-execution isolation
  2. Content from fetched web pages flows into the prompt that writes executed Python, so a hijacking page can leak keys and run arbitrary code unattended. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.05 / 1.00

AgentFlow has no agent identity or authorization layer. Importing the engines and tools loads a .env file into the process environment, and every engine and tool builds its own client from those environment keys. Because model-written Python runs inside the same process, any hijacked step can use every provider key and the OS user's full authority.

C2 Approval gates

Minimal 0.00 / 1.00

There is no approval step anywhere. The only check before running a tool is that the model chose a registered tool name; the command itself is model-written Python that runs straight away. Consequential actions, including arbitrary code on the host and outbound requests, happen with no human involved.

C3 Tool & action scoping

Minimal 0.00 / 1.00

Tool 'commands' are not structured arguments: the executor runs whatever Python the model writes, as long as it ends with an execution = tool.execute(...) line, so any code can be prefixed. The web tools fetch any URL with no host allowlist or internal-address block, and the Python tool's only filter strips exit/quit calls. The default constructor enables every bundled tool, including code execution and web fetch.

C4 Code-execution isolation

Minimal 0.00 / 1.00

Model-generated code runs inside the agent's own Python process twice over: the executor exec()s every tool command with the executor module's globals, and the Python coder tool exec()s generated code (once with no namespace at all). There is no container, subprocess, or even an import filter, and the process holds every API key and full network access.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Web pages, Wikipedia articles, and search results are stored as step results and pasted straight into the planner's and executor's prompts with the same standing as the user's question. The executor then writes Python from that context and runs it. A page that hijacks the model can therefore exfiltrate keys through an arbitrary request and run destructive code, unattended.

C6 Memory, context & configuration integrity

Minimal 0.30 / 1.00

Agent memory lives only for one solve() call, and there are no auto-loaded instruction files. The one persistent store is an opt-in disk cache of LLM responses (off by default in the engine factory) that, when enabled, replays stored responses, including generated tool commands, for any identical prompt in later sessions without validation. The .env file is found relative to the package source, not the working directory (python-dotenv behaviour, inferred).

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

The inference agent loads no third-party code at runtime: tools are imported only from the package's own tools directory, there is no MCP or plugin loader, and LLM engines are API clients. Training code (verl) defaults trust_remote_code to False. This is absence of the surface, not a control.

C8 Secrets & sensitive-data protection

Minimal 0.05 / 1.00

Secrets are plain environment variables loaded from .env at import, with no masking anywhere. Verbose mode, on by default, prints every prompt, command and tool result to stdout. Model-written code runs in the same process and can read every provider key from os.environ.

C9 Audit & traceability

Minimal 0.28 / 1.00

The solver keeps a per-run dictionary of prompts, generated commands and tool results and returns it to the caller, and prints each step to stdout. Nothing is written to a durable or tamper-resistant log, there is no actor attribution, and the record lives in the same process that runs model code.

C10 Limits & kill switch

Minimal 0.40 / 1.00

The solve loop stops after 10 steps or 300 seconds by default, checked between steps, and each tool block waits at most 120 seconds. But timeouts only stop waiting: the daemon thread running model code keeps going, and the Python tool's 10-second timer raises in the wrong thread so it never interrupts anything. There is no token or cost budget, and in-process model code could alter the limits.