# Defense-in-depth score: Windmill

**Repo:** https://github.com/windmill-labs/windmill · **Commit:** `85248eee786b633def40c42b6007fc7b907fb09e` (1.823.1) · **Reviewed:** 2026-10-05
**What it is:** Open-source developer platform for scripts, workflows and internal apps, with AI agent steps that call scripts, flows, MCP servers and web search as tools.
**Category:** Agent Frameworks
**Scored configuration:** Self-hosted docker-compose.yml with the shipped .env (CE image, privileged worker with PID-namespace isolation, nsjail off): an AI agent flow step with default settings (10 turns, compaction memory) and author-attached script, MCP and web-search tools.
**Agent surface (default):** code execution yes · filesystem write yes · network egress yes · external credentials yes · persistent memory yes · untrusted input yes · third party extensions opt-in · sub agents yes · external communication yes

## Score: 3.3 / 10.0 (Minimal)

| # | Criterion | S | C | D | B | Raw | Cap | Score | Confidence |
|---|---|---|---|---|---|---|---|---|---|
| C1 | Identity & least privilege | L2 | L3 | L0 | L1 | 0.42 | G1 | **0.42** (alt) | High |
| C2 | Approval gates | L0 | L0 | L0 | L0 | 0.00 | none | **0.00** | High |
| C3 | Tool & action scoping | L2 | L2 | L3 | L1 | 0.50 | none | **0.50** | High |
| C4 | Code-execution isolation | L2 | L2 | L0 | L1 | 0.35 | G1 | **0.35** (alt) | High |
| C5 | Untrusted input blast radius | L0 | L0 | L0 | L0 | 0.00 | C5-WORSTCASE | **0.00** | High |
| C6 | Memory, context & configuration integrity | L1 | L1 | L2 | L2 | 0.35 | none | **0.35** | High |
| C7 | Third-party extensions | L1 | L1 | L2 | L0 | 0.25 | none | **0.25** | High |
| C8 | Secrets & sensitive-data protection | L2 | L1 | L1 | L0 | 0.28 | none | **0.28** | Medium |
| C9 | Audit & traceability | L3 | L3 | L2 | L2 | 0.65 | none | **0.65** | High |
| C10 | Limits & kill switch | L2 | L2 | L2 | L2 | 0.50 | none | **0.50** | High |


A Windmill AI agent step runs every tool call the model makes without human review, using the full workspace permissions of whoever the flow runs as and the author's long-lived service credentials. If content it reads (a webhook payload, an email, a web page, a tool result) carries injected instructions, it can leak data and act on connected systems unattended. Authorization checks, author-pinned tool inputs, SSRF checks on MCP servers and a working cancel path are real strengths, but in the shipped compose file tool code runs as root in a privileged worker with only process-namespace isolation, and the nsjail sandbox and token scoping are opt-in.

## Critical gaps
- Tool jobs carry the runner's full workspace authority and long-lived service credentials, so a hijacked agent holds everything the runner can reach. (ASI03, T3; C1). Evidence: [backend/windmill-worker/src/ai/tools.rs:518-523](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L518-L523); [backend/windmill-common/src/auth.rs:730](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/auth.rs#L730)
- In the shipped compose file the worker runs privileged as root and jobs get only PID-namespace isolation. (ASI05, T11; C4). Evidence: [docker-compose.yml:78](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/docker-compose.yml#L78); [docker-compose.yml:85](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/docker-compose.yml#L85)
- A prompt-injected agent can both exfiltrate data and take irreversible actions with no human involved. (ASI01, T6, LLM01; C5). Evidence: [backend/windmill-worker/src/ai/tools.rs:122](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L122); [backend/windmill-worker/src/ai/tools.rs:518-523](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L518-L523)
- Third-party Hub scripts run as ordinary worker jobs with the runner's job token and no extra confinement. (ASI04, T17; C7). Evidence: [backend/windmill-worker/src/ai_executor.rs:984](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L984); [backend/windmill-worker/src/bash_executor.rs:281-282](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/bash_executor.rs#L281-L282)

## Criterion details

### C1 Identity & least privilege: 0.42 (high confidence)

Every tool an agent calls runs as a Windmill job under the identity of whoever the flow runs as, with a job token that carries that user's full workspace permissions and lasts as long as the instance's maximum job duration (seven days on a self-hosted install). Authorization is real: resources, secrets and MCP servers are loaded through the job's own permissions, so a flow cannot use what its runner cannot read. Windmill also lets an author cap each step's or tool's token to named API scopes, and a tool can never get a wider token than its agent, but that cap is off unless set. A hijacked agent therefore acts with the runner's whole workspace authority plus the long-lived credentials attached to its tools.

- **default configuration** (default; raw 0.28 → 0.28)
  - **S L1:** Tool jobs inherit the flow runner's identity and an unscoped job token; downstream credentials are the author's long-lived resources. Evidence: [backend/windmill-worker/src/ai/tools.rs:522](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L522); [backend/windmill-worker/src/ai/tools.rs:518-523](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L518-L523); [backend/windmill-common/src/auth.rs:730](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/auth.rs#L730) (verified)
    - *To reach the next level:* No per-tool narrowing by default: read and write tools share the runner's full workspace authority.
  - **C L2:** Windmill tools, nested agents and MCP resource/token loading all go through the job's permissioned client and job permissions. Evidence: [backend/windmill-worker/src/ai/tools.rs:518-523](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L518-L523); [backend/windmill-worker/src/ai/utils.rs:573-574](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/utils.rs#L573-L574); [backend/windmill-types/src/flows.rs:930](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-types/src/flows.rs#L930) (verified)
    - *To reach the next level:* Limited by the strength of the identity: every path gets the same broad authority rather than a per-capability one.
  - **D L1:** Job token scopes default to unset, so tools receive the runner's full permissions; the shipped instance starts with a superadmin account. Evidence: [backend/windmill-common/src/scopes.rs:592](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/scopes.rs#L592); [README.md:209](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/README.md#L209) (verified)
    - *To reach the next level:* Agent tools are not narrowed to minimal scopes by default.
  - **B L0:** A hijacked agent holds the runner's workspace authority (every secret and resource they can read) for up to the job-token lifetime, plus every attached service credential. Evidence: [backend/windmill-common/src/auth.rs:730](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/auth.rs#L730); [backend/windmill-common/src/worker.rs:295](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/worker.rs#L295); [backend/windmill-worker/src/bash_executor.rs:281-282](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/bash_executor.rs#L281-L282) (verified)
    - *To reach the next level:* Authority is not limited to one system or to short-lived credentials.
- **opt-in per-step and per-tool job token scopes (job_token_scopes)** (alt; raw 0.42, cap G1 → 0.42) ← counted
  - **S L2:** Authors can cap a step's or tool's job token to named API scopes, intersected so a tool never exceeds its agent. Evidence: [backend/windmill-common/src/scopes.rs:592](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/scopes.rs#L592); [backend/windmill-types/src/flows.rs:930](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-types/src/flows.rs#L930) (verified)
    - *To reach the next level:* Scopes narrow Windmill API access only; credentials passed to tools as resources are not narrowed.
  - **C L3:** The intersection applies to every Windmill tool job pushed by the agent, including nested agents. Evidence: [backend/windmill-types/src/flows.rs:930](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-types/src/flows.rs#L930); [backend/windmill-types/src/flows.rs:932](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-types/src/flows.rs#L932) (verified)
    - *To reach the next level:* Not evaluated against an end user distinct from the runner; MCP tools use the resource's own token.
  - **D L0:** Off unless the author sets scopes. Evidence: [backend/windmill-common/src/scopes.rs:592](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/scopes.rs#L592) (verified)
    - *To reach the next level:* Off by default.
  - **B L1:** With scopes set, tools still hold the author's long-lived downstream credentials. Evidence: [backend/windmill-common/src/auth.rs:730](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/auth.rs#L730) (verified)
    - *To reach the next level:* Downstream credentials stay long-lived and broadly scoped.
- **Cap:** G1: Opt-in mechanism: off in the scored default configuration.

### C2 Approval gates: 0.00 (high confidence)

An AI agent step runs every tool call the model chooses straight away; there is no approval step for agent tool calls. Windmill flows do have strong human-approval steps, but they sit between flow steps and cannot be attached to an individual tool call inside an agent. Whatever the attached tools can do, including writes, messages and payments through author scripts or MCP servers, happens without a person seeing the exact call.

- **S L0:** No approval in the agent loop or tool executor. Evidence: searched `rg -n -i 'approv|suspend|human'` in `backend/windmill-worker/src/ai_executor.rs backend/windmill-worker/src/ai` → 0 hits (no approval, suspend or human-review step anywhere in the agent loop or tool executor) (verified)
  - *To reach the next level:* No per-call human approval for agent tools.
- **C L0:** No gate exists, so every tool path, including MCP tools, is ungated. Evidence: searched `rg -n -i 'approv|suspend|human'` in `backend/windmill-worker/src/ai_executor.rs backend/windmill-worker/src/ai` → 0 hits (no approval, suspend or human-review step anywhere in the agent loop or tool executor); [backend/windmill-types/src/flows.rs:955-962](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-types/src/flows.rs#L955-L962) (verified)
  - *To reach the next level:* No gate for any tool path.
- **D L0:** Nothing to enable: agent tools carry no approval setting. Evidence: [backend/windmill-types/src/flows.rs:955-962](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-types/src/flows.rs#L955-L962) (verified)
  - *To reach the next level:* No approval on by default.
- **B L0:** Author tools can take irreversible external actions with no preview or undo. Evidence: [backend/windmill-worker/src/ai/utils.rs:87-88](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/utils.rs#L87-L88); [backend/windmill-worker/src/ai/tools.rs:518-523](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L518-L523) (verified)
  - *To reach the next level:* No previews, dry-runs or rollback for consequential tool calls.
- **Cap:** none

### C3 Tool & action scoping: 0.50 (high confidence)

An agent only gets the tools its author attached and cannot load new ones; a run can narrow the roster but never widen it. Inputs the author pinned are removed from what the model sees and always override the model's values, MCP tools can be limited by include and exclude lists, and MCP server addresses are checked against private and metadata addresses without following redirects. Unpinned inputs are passed to the author's code as the model wrote them, and tools are general scripts that reach whatever their credentials reach.

- **S L2:** Typed tool schemas with author-pinned inputs hidden from the model and overriding its values; SSRF-checked MCP URLs. Evidence: [backend/windmill-worker/src/ai_executor.rs:1066](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L1066); [backend/windmill-worker/src/ai/utils.rs:87-88](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/utils.rs#L87-L88); [backend/windmill-mcp/src/client/mod.rs:46](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-mcp/src/client/mod.rs#L46); [backend/windmill-mcp/src/client/mod.rs:86](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-mcp/src/client/mod.rs#L86) (verified)
  - *To reach the next level:* Unpinned arguments are not checked against allowlists or bounds before reaching tool code.
- **C L2:** Pinned-input handling applies to every Windmill tool; MCP tools get name filters but arguments pass through unchanged. Evidence: [backend/windmill-worker/src/ai/utils.rs:87-88](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/utils.rs#L87-L88); [backend/windmill-worker/src/ai/utils.rs:632](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/utils.rs#L632) (verified)
  - *To reach the next level:* MCP tool arguments and model-chosen values are not validated by a shared layer.
- **D L3:** No tool is attached by default; the author adds each one and a run may only narrow the set. Evidence: [backend/windmill-worker/src/ai_executor.rs:889](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L889); [backend/windmill-worker/src/ai_executor.rs:943](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L943) (verified)
  - *To reach the next level:* No per-task least-privilege tool sets beyond what the author attaches.
- **B L1:** Tools are general-purpose scripts and MCP servers acting on production systems with the author's credentials. Evidence: [backend/windmill-worker/src/ai/tools.rs:518-523](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L518-L523); [backend/windmill-common/src/auth.rs:730](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/auth.rs#L730) (verified)
  - *To reach the next level:* Tool reach is not bounded by quantity or scope limits.
- **Cap:** none

### C4 Code-execution isolation: 0.35 (high confidence)

Agent tools run as Windmill jobs on workers. In the shipped Docker Compose setup the only job isolation is a separate process namespace, and the worker container itself runs privileged and as root, so a tool job can reach the worker host and its network. Windmill ships an nsjail sandbox that limits the filesystem and drops capabilities, but it is off by default and leaves network access and the job's token in place. Jobs start with a cleared environment that adds only Windmill's own variables, including the job token.

- **default configuration** (default; raw 0.33, cap C4-HOSTROOT → 0.25)
  - **S L1:** Default isolation is a PID namespace (unshare) with a per-job directory; nsjail is disabled unless configured. Evidence: [docker-compose.yml:85](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/docker-compose.yml#L85); [backend/windmill-worker/src/worker.rs:370-373](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/worker.rs#L370-L373) (verified)
    - *To reach the next level:* No filesystem, user or network separation by default.
  - **C L2:** The same isolation wrapper applies to every language executor that agent tools use. Evidence: [backend/windmill-worker/src/bash_executor.rs:281-282](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/bash_executor.rs#L281-L282); [backend/windmill-worker/src/worker.rs:1010-1017](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/worker.rs#L1010-L1017) (verified)
    - *To reach the next level:* Coverage of a weak primitive; no stronger boundary on any path by default.
  - **D L2:** PID isolation is turned on by the shipped compose file and can be switched off by an env var or instance setting. Evidence: [docker-compose.yml:85](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/docker-compose.yml#L85); [backend/windmill-worker/src/worker.rs:1010-1017](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/worker.rs#L1010-L1017) (verified)
    - *To reach the next level:* Isolation is not fail-closed and not protected from silent operator changes.
  - **B L0:** Jobs run as root inside a privileged worker container with full network and the job token in the environment. Evidence: [docker-compose.yml:78](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/docker-compose.yml#L78); searched `rg -n '^USER'` in `Dockerfile` → 0 hits (the runtime image never switches away from root); [backend/windmill-worker/src/bash_executor.rs:281-282](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/bash_executor.rs#L281-L282) (verified)
    - *To reach the next level:* Worker runs privileged as root; L1+ needs an unprivileged container without host-equivalent rights.
- **opt-in nsjail sandbox (DISABLE_NSJAIL=false or job_isolation=nsjail)** (alt; raw 0.35, cap G1 → 0.35) ← counted
  - **S L2:** nsjail with bind-mounted system directories, dropped capabilities and rlimits, but network namespaces are not used. Evidence: [backend/windmill-worker/nsjail/run.python3.config.proto:15](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/nsjail/run.python3.config.proto#L15); [backend/windmill-worker/nsjail/run.python3.config.proto:19](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/nsjail/run.python3.config.proto#L19) (verified)
    - *To reach the next level:* Network is not denied by default inside the jail.
  - **C L2:** Per-language nsjail configs exist for the script runtimes agent tools use. Evidence: [backend/windmill-worker/nsjail/run.python3.config.proto:1-3](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/nsjail/run.python3.config.proto#L1-L3); [backend/windmill-worker/src/bash_executor.rs:281-282](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/bash_executor.rs#L281-L282) (verified)
    - *To reach the next level:* Not shown to cover every job kind without exception.
  - **D L0:** Off by default. Evidence: [backend/windmill-worker/src/worker.rs:370-373](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/worker.rs#L370-L373); [docker-compose.yml:197](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/docker-compose.yml#L197) (verified)
    - *To reach the next level:* Off by default.
  - **B L1:** Inside the jail jobs keep full network egress and the job token. Evidence: [backend/windmill-worker/nsjail/run.python3.config.proto:15](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/nsjail/run.python3.config.proto#L15); [backend/windmill-worker/nsjail/run.python3.config.proto:20](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/nsjail/run.python3.config.proto#L20) (verified)
    - *To reach the next level:* Egress and credentials remain available inside the sandbox.
- **Cap:** G1: Opt-in mechanism: off in the scored default configuration.

### C5 Untrusted input blast radius: 0.00 (high confidence)

Agent steps commonly read content their author did not write: flow inputs from webhooks, email and other triggers, web search results, and the output of tools and MCP servers, which all enter the conversation like any other tool result. Nothing marks that content as untrusted or restricts what the agent may do after reading it, and no tool call needs approval. A successful prompt injection can therefore use any attached tool to send data out and take irreversible actions with the runner's authority.

- **S L0:** No structural limit after untrusted content is read. Evidence: searched `rg -n -i 'untrusted|prompt.injection|guardrail|taint|provenance'` in `backend/windmill-worker/src/ai_executor.rs backend/windmill-worker/src/ai` → 0 hits (no provenance marking, detection or taint tracking for tool results or web content) (verified)
  - *To reach the next level:* No approval or capability restriction triggered by untrusted content.
- **C L0:** Tool, MCP and web-search results enter context as ordinary tool messages. Evidence: [backend/windmill-worker/src/ai/tools.rs:122](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L122); [backend/windmill-worker/src/ai_executor.rs:943](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L943) (verified)
  - *To reach the next level:* Untrusted sources are not distinguished.
- **D L0:** No control exists to be on. Evidence: searched `rg -n -i 'untrusted|prompt.injection|guardrail|taint|provenance'` in `backend/windmill-worker/src/ai_executor.rs backend/windmill-worker/src/ai` → 0 hits (no provenance marking, detection or taint tracking for tool results or web content) (verified)
  - *To reach the next level:* Nothing on by default.
- **B L0:** A hijacked agent can leak secrets and data through any egress tool and take irreversible actions, unattended. Evidence: [backend/windmill-worker/src/ai/utils.rs:87-88](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/utils.rs#L87-L88); [backend/windmill-worker/src/ai/tools.rs:518-523](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L518-L523); [backend/windmill-common/src/auth.rs:730](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/auth.rs#L730) (verified)
  - *To reach the next level:* Exfiltration and irreversible actions are not gated.
- **Cap:** C5-WORSTCASE: Worst case (B L0): a hijacked agent can leak data and take irreversible actions unattended.

### C6 Memory, context & configuration integrity: 0.35 (high confidence)

Agent steps keep conversation memory by default, summarising older turns when the context fills up and storing the whole conversation, including tool results, in the database for the next run of the same conversation. Memory is kept per workspace, flow and conversation id, and the model cannot choose which memory it reads. Nothing validates what is written, so injected text that lands in a tool result or summary is replayed in every later turn of that conversation. There are no workspace instruction files or repository configs that an agent loads.

- **S L1:** Full conversation (tool results and compaction summaries) is persisted and replayed without validation. Evidence: [backend/windmill-worker/src/ai_executor.rs:2375](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L2375); [backend/windmill-worker/src/ai_executor.rs:1407](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L1407); [frontend/src/lib/components/flows/agentFormFields.ts:58](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/frontend/src/lib/components/flows/agentFormFields.ts#L58) (verified)
  - *To reach the next level:* Memory writes are not gated, validated or provenance-tagged.
- **C L1:** The one memory store has no write control; summaries and tool results persist as is. Evidence: [backend/windmill-worker/src/memory_common.rs:79](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/memory_common.rs#L79) (verified)
  - *To reach the next level:* No control over summaries or tool-result content entering memory.
- **D L2:** Memory is keyed by workspace, flow and conversation id in queries, and the model cannot pick the id. Evidence: [backend/windmill-worker/src/ai_executor.rs:349](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L349); [backend/windmill-worker/src/ai_executor.rs:251](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L251) (verified)
  - *To reach the next level:* Nothing stops a flow from using a shared or caller-chosen conversation id; no retention limit by default.
- **B L2:** Poisoned memory persists for the conversation and can steer later tool calls, but stays inside one conversation. Evidence: [backend/windmill-worker/src/ai_executor.rs:1407](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L1407); [backend/windmill-worker/src/memory_common.rs:79](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/memory_common.rs#L79) (verified)
  - *To reach the next level:* Poisoned memory can still trigger ungated tool use in later turns.
- **Cap:** none

### C7 Third-party extensions: 0.25 (high confidence)

Agents can use third-party code in two ways: Hub scripts, which are pinned to a version id and fetched from Windmill's hub, and remote MCP servers, whose tool lists are fetched fresh on every run with no pinning or change detection. Nothing third-party is enabled by default; the author adds each one. MCP servers run remotely and only receive their own token and the call arguments, but Hub scripts run on the worker like any other job, with the runner's job token and the same weak default isolation.

- **S L1:** MCP tool definitions are fetched each run without pinning; Hub scripts are pinned by version id without an integrity check. Evidence: [backend/windmill-worker/src/ai_executor.rs:1101](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L1101); [backend/windmill-common/src/scripts.rs:266-268](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/scripts.rs#L266-L268); searched `rg -n -i 'sha256|checksum|integrity|signature'` in `backend/windmill-mcp/src/client` → 0 hits (no pinning or integrity check of MCP server tool definitions) (verified)
  - *To reach the next level:* No pinning or integrity check for MCP servers, and no re-approval when tools change.
- **C L1:** Version pinning covers Hub scripts only. Evidence: [backend/windmill-worker/src/ai_executor.rs:984](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L984); [backend/windmill-common/src/scripts.rs:266-268](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/scripts.rs#L266-L268) (verified)
  - *To reach the next level:* MCP servers are not verified.
- **D L2:** Nothing third-party is enabled by default; the author explicitly adds a Hub script or MCP resource. Evidence: [backend/windmill-worker/src/ai/utils.rs:573-574](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/utils.rs#L573-L574); [backend/windmill-worker/src/ai_executor.rs:984](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L984) (verified)
  - *To reach the next level:* Adding an MCP server does not show or approve the exact tool set that will run.
- **B L0:** Hub scripts run as worker jobs with the runner's job token in a privileged, root worker. Evidence: [backend/windmill-worker/src/bash_executor.rs:281-282](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/bash_executor.rs#L281-L282); [docker-compose.yml:78](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/docker-compose.yml#L78) (verified)
  - *To reach the next level:* Third-party scripts are not confined or given scoped credentials.
- **Cap:** none

### C8 Secrets & sensitive-data protection: 0.28 (medium confidence)

Secret variables are encrypted at rest with a per-workspace key, and any secret a job fetches, along with its own token, is masked in that job's logs. Jobs start from a cleared environment. Author-pinned credentials are kept out of the tool schema the model sees. Redaction covers job logs only: job results, stored conversation memory and messages sent to the model provider are not redacted, instance-level settings are stored unencrypted under the default backend, and every tool job receives a long-lived token with the runner's full permissions. Usage statistics code is closed source in this repository.

- **S L2:** Encrypted secret variables and per-job log masking of fetched secrets and the job token. Evidence: [backend/windmill-common/src/secret_backend/database.rs:11](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/secret_backend/database.rs#L11); [backend/windmill-common/src/sensitive_log_masks.rs:244](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/sensitive_log_masks.rs#L244); [backend/windmill-worker/src/worker.rs:4536-4541](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/worker.rs#L4536-L4541) (verified)
  - *To reach the next level:* No redaction before model-bound messages, results or memory; instance settings stored in plaintext.
- **C L1:** Masking covers job logs; subprocess environments are cleared apart from Windmill's own variables. Evidence: [backend/windmill-worker/src/worker.rs:4536-4541](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/worker.rs#L4536-L4541); searched `rg -n -i 'redact|mask'` in `backend/windmill-worker/src/ai_executor.rs backend/windmill-worker/src/ai` → 2 hits (both hits are comments about job-log masks; nothing redacts tool results or model-bound messages) (verified)
  - *To reach the next level:* Results, transcripts, memory and model-bound messages are not covered.
- **D L1:** Log masking is always on; usage statistics are sent by closed-source code whose content cannot be checked here. Evidence: [backend/windmill-common/src/stats_oss.rs:12-15](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/stats_oss.rs#L12-L15) (inferred)
  - *To reach the next level:* Telemetry is not shown to be opt-in or content-free.
- **B L0:** Every tool subprocess receives a job token with the runner's full workspace authority, valid for up to the maximum job duration. Evidence: [backend/windmill-worker/src/bash_executor.rs:281-282](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/bash_executor.rs#L281-L282); [backend/windmill-common/src/auth.rs:730](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/auth.rs#L730); [backend/windmill-common/src/worker.rs:295](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/worker.rs#L295) (verified)
  - *To reach the next level:* Tokens reachable by tool code are long-lived and broadly privileged.
- **Cap:** none

### C9 Audit & traceability: 0.65 (high confidence)

Each tool call is recorded as it happens in the flow's status, and Windmill tools run as separate jobs whose arguments, results, logs, timestamps, creator and run-as identity are stored with parent and root job ids, so a run can be traced across nested agents. MCP calls are recorded with their arguments, and the agent's final result carries the full message history. The audit log that would record configuration changes and secret reads is an Enterprise feature and does nothing in this code. Records live in the instance database, which tool jobs can reach through the API with the runner's permissions.

- **S L3:** Structured job records with arguments and results plus creator and run-as identity and parent/root correlation ids. Evidence: [backend/windmill-worker/src/ai/tools.rs:522](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L522); [backend/windmill-worker/src/ai/tools.rs:518-523](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L518-L523); [backend/windmill-worker/src/ai/tools.rs:342](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L342) (verified)
  - *To reach the next level:* Records are not tamper-evident or exported to an external store by default.
- **C L3:** Windmill tool calls, MCP calls and nested agents are all recorded. Evidence: [backend/windmill-worker/src/ai/tools.rs:342](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L342); [backend/windmill-worker/src/ai/tools.rs:210](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L210) (verified)
  - *To reach the next level:* Configuration changes, memory writes and credential use are not recorded in the open-source build.
- **D L2:** On by default and written by the worker, but stored where the runner's token can reach it. Evidence: [backend/windmill-audit/src/audit_oss.rs:87](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-audit/src/audit_oss.rs#L87); [backend/windmill-common/src/auth.rs:730](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/auth.rs#L730) (verified)
  - *To reach the next level:* The record is not written by a component the agent's tools cannot alter.
- **B L2:** Tool-call actions are flushed as each call starts; full message history lands when the step ends. Evidence: [backend/windmill-worker/src/ai/tools.rs:342](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L342); [backend/windmill-worker/src/ai/tools.rs:210](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai/tools.rs#L210) (verified)
  - *To reach the next level:* No fail-closed recording; results and messages are only complete at the end of the step.
- **Cap:** none

### C10 Limits & kill switch: 0.50 (high confidence)

An agent step stops after 10 model turns by default, an author can raise that to at most 1,000, and every model request and tool job is bounded by the job timeout, which defaults to seven days on a self-hosted install. Cancelling a run stops the loop, gives in-flight tools 30 seconds and then aborts them and cancels any tool jobs still queued or running. There is no token or cost budget, and a nested agent tool starts its own turn budget.

- **S L2:** Iteration cap plus per-request and per-job timeouts, with a cancel path that aborts in-flight tools after a grace period. Evidence: [backend/windmill-worker/src/ai_executor.rs:104](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L104); [backend/windmill-worker/src/ai_executor.rs:1586](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L1586); [backend/windmill-worker/src/ai_executor.rs:1143](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L1143); [backend/windmill-worker/src/ai_executor.rs:1234](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L1234) (verified)
  - *To reach the next level:* No token or cost budget and no repeated-action breaker.
- **C L2:** Limits cover the loop and tool jobs (job timeouts); nested agents get their own turn budget. Evidence: [backend/windmill-worker/src/ai_executor.rs:688](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L688); [backend/windmill-worker/src/ai_executor.rs:1635](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L1635) (verified)
  - *To reach the next level:* Nested agents do not count against the parent's budget.
- **D L2:** A sensible turn default with a hard ceiling the model cannot raise, but a seven-day default time ceiling. Evidence: [backend/windmill-worker/src/ai_executor.rs:104](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L104); [backend/windmill-worker/src/ai_executor.rs:1586](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L1586); [backend/windmill-common/src/worker.rs:295](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/worker.rs#L295) (verified)
  - *To reach the next level:* Delegation to a nested agent resets the turn budget.
- **B L2:** Moderate turn ceiling; stopping cancels pending tool jobs after 30 seconds, but run time can reach days. Evidence: [backend/windmill-worker/src/ai_executor.rs:1143](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L1143); [backend/windmill-worker/src/ai_executor.rs:1234](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-worker/src/ai_executor.rs#L1234); [backend/windmill-common/src/worker.rs:295](https://github.com/windmill-labs/windmill/blob/85248eee786b633def40c42b6007fc7b907fb09e/backend/windmill-common/src/worker.rs#L295) (verified)
  - *To reach the next level:* No tight per-run time or spend ceiling.
- **Cap:** none

## Rule-of-Two check
[A] untrusted input: Flow inputs from webhooks, email and other triggers, provider web search (backend/windmill-worker/src/ai_executor.rs:943), MCP and tool results as tool messages (backend/windmill-worker/src/ai/tools.rs:122) · [B] sensitive data/systems: Job token with the runner's workspace permissions (backend/windmill-common/src/auth.rs:730) and author-attached resources (backend/windmill-worker/src/ai/utils.rs:87) · [C] state change / egress: Author scripts, flows and MCP tools pushed as jobs with no gate (backend/windmill-worker/src/ai/tools.rs:518) · Same default session? Yes

## Highest-impact improvements
1. Offer per-tool human approval for agent tool calls that shows the exact arguments, reusing the flow approval machinery. (C2 S L0→L3, +0.225 before caps; Playbook 5)
2. Once untrusted content enters an agent run, route egress and state-changing tools through approval. (C5 S L0→L2, +0.150 before caps; Playbook 1)
3. Ship the compose worker unprivileged and non-root, and make nsjail with network isolation the default. (C4 B L0→L2, +0.100 before caps; Playbook 3)
4. Default agent tool job tokens to a minimal scope set and shorten their lifetime to the tool's own timeout. (C1 D L1→L2, +0.050 before caps; Playbook 4)
5. Add a run-level token or cost budget shared with nested agent tools. (C10 S L2→L3, +0.075 before caps; Playbook 3 step 3)

## Re-audit log
- C1 S: L2 → L1. Workspace RBAC scopes the user, but tool jobs get the runner's whole authority and unscoped long-lived downstream resources; between L1 and L2, so the lower level.
- C6 B: L3 → L2. Conversation memory is session-like, but the run's memory id can come from the caller, so persistence is not provably limited to one session; poisoned memory can still drive ungated tool calls.

## Limitations
- Static source review of the pinned commit only; nothing was executed, installed, or probed.
- The version is taken from version.txt; release tags were not available in the shallow clone.
- Enterprise-only code (audit logs, memory_ee, stats_ee and other *_ee modules) is not in this repository; the CE Docker image may contain closed-source code that behaves differently from the open-source stubs reviewed here.
- Scored the AI agent flow step; the editor AI assistant, the Windmill MCP server endpoint and the AI proxy were not scored as agents and were only read where they affect the agent step.
- Frontend rendering of agent output and the Helm chart (not in this repository) were not examined.
