BoundBench

TaskWeaver

Microsoft's code-first agent framework that plans data-analytics tasks and executes LLM-generated Python in a stateful Jupyter kernel.

github.com/microsoft/taskweaver · 2026-10-04 · d44ddef

Defense-in-depth score

3.0 / 10

Minimal

TaskWeaver runs every piece of model-written Python immediately, with no human approval, so its safety rests almost entirely on the per-session Docker container it uses by default. That container keeps host credentials out and fails closed when Docker is missing, but it is unhardened, has full network egress and no resource limits. A prompt injection in data or tool output can therefore exfiltrate the user's files or call external APIs unattended.

Key gaps (1)

  1. A hijacked TaskWeaver session can read the user's data files and send them anywhere, or call arbitrary external APIs, with no human in the loop: model code executes immediately in a container with full network egress. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.47 / 1.00

TaskWeaver has no per-request authorization layer or scoped identity of its own. In the default container mode the code kernel gets a scrubbed environment with only TaskWeaver variables, so the LLM API key held by the host process is not visible to generated code; but every plugin's configuration is loaded into the same kernel the model's code runs in. Switching to local mode, a single config value, hands the kernel a full copy of the host environment, which the code itself marks as a TODO to filter.

C2 Approval gates

Minimal 0.05 / 1.00

There is no human approval step anywhere in TaskWeaver. Code generated by the model is optionally checked by a static filter (off by default in container mode) and then executed immediately; plugin calls happen inside that code. A search of the Python source finds no confirmation or approval logic, and the README's own wish-list asks for 'user confirmation before running the plugin'.

C3 Tool & action scoping

Minimal 0.05 / 1.00

The core 'tool' is arbitrary Python executed in a Jupyter kernel, so a hijacked agent can do anything the kernel can. An AST-based module allowlist and function blocklist exists, but it is off by default in the default container mode and is credited under code-execution isolation, not here. Plugins declare typed parameters in YAML, but the model calls them from free-form code, so nothing enforces those schemas.

C4 Code-execution isolation

Moderate 0.57 / 1.00

By default every piece of model-written code, including CLI-only shell commands and plugin code, runs in a per-session Docker container that is started from a stock python:3.10-slim based image. The container gets a scrubbed environment and only the session's working and kernel directories mounted, and TaskWeaver refuses to start rather than falling back to the host if Docker is unavailable. But the container is not hardened: no dropped capabilities, no resource limits, default network with full egress, kernel ports published on the host, and its entrypoint runs as root before dropping to a user with the host UID. Local mode, a single config value, runs code on the host behind an escapable AST filter.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Nothing in TaskWeaver separates untrusted content from instructions. Execution output, plugin results (the default-enabled Klarna search plugin fetches web data), and the content of any file the code reads are appended to the planner's conversation as ordinary user-role messages. Because generated code runs without approval in a container with full network egress, an injected instruction can both leak the user's data and take irreversible external actions unattended.

C6 Memory, context & configuration integrity

Moderate 0.55 / 1.00

In the default configuration conversation memory lives only in the running session, and the long-term 'experience' feature is off; when enabled, experiences are only created when the user runs /save, but the saved chat (including any injected content) is LLM-summarised and re-injected into every later session's prompt for that project, with no provenance or expiry. Config, plugins and examples load from the project directory, which the README has the operator pass explicitly; when -p is omitted TaskWeaver walks up from the current directory to find a taskweaver_config.json. In the default container mode generated code cannot write the project directory.

C7 Third-party extensions

Minimal 0.17 / 1.00

Plugins are Python files placed in the project's plugins folder; any YAML there with enabled: true is picked up by glob and its code executed in the same kernel as model code, with no pinning, hash or consent prompt. Two shipped plugins are enabled by default. The sandbox image itself is pulled as an unpinned ':latest' tag from Docker Hub and silently re-pulled whenever the registry copy changes. Generated code can also pip-install packages inside the container because network is open.

C8 Secrets & sensitive-data protection

Minimal 0.25 / 1.00

The LLM API key is read from the project's taskweaver_config.json in plaintext or from environment variables, and nothing in the codebase masks or redacts secrets. The default container mode keeps host environment variables out of the code kernel, which is the one protected path. Every LLM prompt is dumped to a JSON file in the session directory by default and the kernel writes a DEBUG-level log; remote telemetry and tracing are opt-in.

C9 Audit & traceability

Minimal 0.42 / 1.00

TaskWeaver writes a plain-text log in the project's logs folder that records every piece of code before it runs and every message between planner and workers, plus a JSON dump of every LLM prompt in the session directory. In the default container mode these files sit outside what generated code can reach. There is no structured per-action record with result status, no actor or approver attribution, and OpenTelemetry tracing is opt-in.

C10 Limits & kill switch

Minimal 0.45 / 1.00

Each user request is capped at 10 internal planner/worker exchanges and three retries per code attempt, and the host stops waiting for a code execution after 180 seconds of silence. There is no token, cost or wall-clock budget, and the timeout does not interrupt the kernel, as a TODO in the code admits, so long-running code keeps going until the session is stopped and its container removed. The container has no CPU or memory limits.