# Defense-in-Depth Score: AutoGPT Platform

**Repo:** https://github.com/Significant-Gravitas/AutoGPT · **Commit:** `f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588` · **Reviewed:** 2026-10-03
**What it is:** Platform for building, deploying and running continuous agents
**Category:** Agent Frameworks
**Scored configuration:** Self-hosted autogpt_platform via the shipped docker-compose with `make init-env` defaults: no LaunchDarkly key (all default-off flags off), no E2B key (bubblewrap shell), AutoPilot chat plus builder graphs.
**Agent surface (default):** code execution yes · filesystem write yes · network egress yes · external credentials yes · persistent memory yes · untrusted input yes · third party extensions yes · sub agents yes · external communication yes

## Score: 4.2 / 10.0 (Minimal)

| # | Criterion | S | C | D | B | Raw | Cap | Score | Confidence |
|---|---|---|---|---|---|---|---|---|---|
| C1 | Identity & least privilege | L2 | L3 | L2 | L1 | 0.53 | — | **0.53** | High |
| C2 | Approval gates | L3 | L0 | L1 | L0 | 0.28 | C2-POWERBYPASS | **0.25** | High |
| C3 | Tool & action scoping | L2 | L2 | L1 | L1 | 0.40 | — | **0.40** | High |
| C4 | Code-execution isolation | L2 | L2 | L3 | L3 | 0.60 | — | **0.60** | High |
| C5 | Untrusted input blast radius | L1 | L2 | L0 | L0 | 0.23 | G1 | **0.23** (alt) | High |
| C6 | Memory, context & configuration integrity | L1 | L1 | L2 | L1 | 0.30 | — | **0.30** | High |
| C7 | Third-party extensions | L1 | L1 | L0 | L3 | 0.30 | — | **0.30** | High |
| C8 | Secrets & sensitive-data protection | L2 | L2 | L3 | L1 | 0.50 | — | **0.50** | High |
| C9 | Audit & traceability | L3 | L2 | L3 | L2 | 0.62 | — | **0.62** | High |
| C10 | Limits & kill switch | L2 | L2 | L2 | L1 | 0.45 | — | **0.45** | High |


AutoGPT agents act with every account a user connects and can browse, fetch, run a sandboxed shell, call any block and run on schedules or webhooks. Code execution is well contained by default (bubblewrap with no network), credentials are encrypted and kept out of the model, and AutoPilot pauses before flagged irreversible blocks. The dominant risk is prompt injection: nothing stops a hijacked agent from sending data out through web fetch or the HTTP block and acting through unflagged blocks, and builder graphs never pause by default. The stronger approval modes and the injection judge ship behind a feature flag that is off on a self-hosted install.

## Critical gaps
- The generic HTTP request block, catalogue MCP writes and browser actions run with no human approval by default, so the approval gate is bypassed by the most general outward action path. (ASI02, ASI09, T10; C2) — [autogpt_platform/backend/backend/blocks/http.py:71](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/http.py#L71); [autogpt_platform/backend/backend/copilot/capabilities/mcp_review.py:55-56](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/capabilities/mcp_review.py#L55-L56)
- A prompt-injected session can exfiltrate data and take irreversible actions unattended: no default control limits egress or state change after reading untrusted content, and builder graphs triggered by webhooks never pause. (ASI01, LLM01, T6; C5) — [autogpt_platform/backend/backend/copilot/tools/web_fetch.py:208-209](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/web_fetch.py#L208-L209); [autogpt_platform/backend/backend/blocks/http.py:71](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/http.py#L71); [autogpt_platform/backend/backend/api/features/library/db.py:570](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/api/features/library/db.py#L570); [autogpt_platform/backend/backend/copilot/gate/__init__.py:101](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/__init__.py#L101)

## Criterion details

### C1 Identity & least privilege — 0.53 (high)

Agents act with the credentials each user has connected (OAuth grants and API keys), and every block execution looks those credentials up by the requesting user's id, failing if the credential is not theirs. OAuth logins request the scopes the blocks ask for, but one credential serves both reads and writes for the whole run, and nothing issues short-lived or per-task tokens. When the optional E2B sandbox is enabled, the AutoPilot shell receives a token for every connected provider. A hijacked agent therefore holds write access across all of a user's connected services.

- **S L2:** Per-user OAuth credentials scoped to the scopes blocks request, static for the run, with one credential for reads and writes. — [autogpt_platform/backend/backend/api/features/integrations/router.py:154](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/api/features/integrations/router.py#L154); [autogpt_platform/backend/backend/integrations/creds_manager.py:146](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/integrations/creds_manager.py#L146) (verified)
  - *To reach the next level:* Read and write paths do not use separate, narrower credentials, and no per-request token downscoping exists.
- **C L3:** Graph and AutoPilot block execution resolves credentials through the user-scoped credential manager and fails when the id does not belong to the requesting user. — [autogpt_platform/backend/backend/executor/manager.py:403-416](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/executor/manager.py#L403-L416); [autogpt_platform/backend/backend/integrations/creds_manager.py:146](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/integrations/creds_manager.py#L146); [autogpt_platform/backend/backend/copilot/integration_creds.py:342](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/integration_creds.py#L342); [autogpt_platform/backend/backend/copilot/tools/bash_exec.py:215](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/bash_exec.py#L215) (verified)
  - *To reach the next level:* The opt-in E2B shell injects every connected provider's token into the sandbox env instead of passing a per-request authorization check, so not every path is authorized per request.
- **D L2:** Default authority is whatever scopes the user grants when connecting an integration; widening is a new OAuth consent, with no default read-only posture. — [autogpt_platform/backend/backend/api/features/integrations/router.py:154](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/api/features/integrations/router.py#L154) (verified)
  - *To reach the next level:* No read-only default identity; write scopes are granted at connect time rather than through explicit time-bounded elevation.
- **B L1:** A hijacked agent can write across every service the user connected (GitHub, Google, Slack, email, payments blocks). — [autogpt_platform/backend/backend/blocks/email_block.py:101](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/email_block.py#L101); [autogpt_platform/backend/backend/blocks/rmfg/pay_cart.py:79](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/rmfg/pay_cart.py#L79); [autogpt_platform/backend/backend/executor/manager.py:403-416](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/executor/manager.py#L403-L416) (verified)
  - *To reach the next level:* Credentials are long-lived and not narrowed to one system or to read-mostly operations.
- **Cap:** none
- **Notes:** Expert credential grants (grant_expert_credential) exist but sit behind the hire-experts flag, which is off without LaunchDarkly. Operator-level platform API keys are also offered to users' blocks; they never reach model context.

### C2 Approval gates — 0.25 (high)

By default AutoPilot pauses before running a block that is flagged irreversible (for example send email, pay cart, post to Slack) and shows the user the exact input, which they can edit, approve or reject. But the flag covers only some blocks: the generic HTTP request block, catalogue MCP tools, web fetch, browser actions and the shell all run without asking. Graphs built in the visual builder default to never pausing. The stronger per-chat approval system (Ask First / Auto modes) is shipped behind a feature flag that is off on a self-hosted install, and in its own default Auto mode it lets an LLM judge decide shell and platform actions.

- **default configuration** (default; raw 0.28, cap C2-POWERBYPASS → 0.25) ← counted
  - **S L3:** Flagged irreversible blocks pause for a per-call human review that stores the exact input payload, lets the reviewer edit or reject it, and runs the reviewed data. — [autogpt_platform/backend/backend/copilot/tools/helpers.py:1054](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/helpers.py#L1054); [autogpt_platform/backend/backend/blocks/_base.py:857](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/_base.py#L857) (verified)
    - *To reach the next level:* No argument-level allow/deny policy; risk tiering is a static per-block boolean.
  - **C L0:** The most powerful general action, the generic HTTP request block, is never flagged and runs ungated, as do catalogue MCP writes and browser actions. — [autogpt_platform/backend/backend/blocks/http.py:71](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/http.py#L71); searched `rg -n 'is_irreversible_action'` in `autogpt_platform/backend/backend/blocks/http.py` → 0 hits (The generic HTTP request blocks never set is_irreversible_action, so the default review never fires for them.); [autogpt_platform/backend/backend/copilot/capabilities/mcp_review.py:55-56](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/capabilities/mcp_review.py#L55-L56) (verified)
    - *To reach the next level:* Generic HTTP, browser actions and catalogue MCP writes would need to pass the same review.
  - **D L1:** Review is on by default for AutoPilot block calls, but graphs created in the builder default sensitive_action_safe_mode to False, so their scheduled and triggered runs never pause. — [autogpt_platform/backend/backend/api/features/library/db.py:570](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/api/features/library/db.py#L570); [autogpt_platform/backend/backend/data/graph.py:72-74](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/data/graph.py#L72-L74); [autogpt_platform/backend/backend/api/features/graphs/routes.py:125](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/api/features/graphs/routes.py#L125); [autogpt_platform/backend/backend/copilot/model.py:163-166](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/model.py#L163-L166) (verified)
    - *To reach the next level:* Builder-created graphs (the continuous-agent path) should default to pausing irreversible actions.
  - **B L0:** A wrongly approved or ungated action can send email, post publicly, or pay with no undo. — [autogpt_platform/backend/backend/blocks/email_block.py:101](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/email_block.py#L101); [autogpt_platform/backend/backend/blocks/rmfg/pay_cart.py:79](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/rmfg/pay_cart.py#L79) (verified)
    - *To reach the next level:* No rollback, preview or quantity bounds on external actions.
- **opt-in AutoPilot approval modes (copilot-auto-mode flag)** (alt; raw 0.23, cap G1 → 0.23)
  - **S L1:** In the gate's default Auto mode, shell and platform actions are decided by an LLM supervisor that can only escalate to a question; only external effects always ask. — [autogpt_platform/backend/backend/copilot/gate/policy.py:173-176](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/policy.py#L173-L176) (verified)
    - *To reach the next level:* Shell and platform calls would need per-call human approval rather than an LLM judge.
  - **C L2:** Every registry tool funnels through the gate, but non-interactive (automation) sessions and chat-bot platforms are ungated and web_fetch counts as a read. — [autogpt_platform/backend/backend/copilot/gate/__init__.py:97](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/__init__.py#L97); [autogpt_platform/backend/backend/copilot/gate/__init__.py:63](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/__init__.py#L63); [autogpt_platform/backend/backend/copilot/gate/policy.py:73](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/policy.py#L73) (verified)
    - *To reach the next level:* Automation sessions, bot platforms and egress-capable reads bypass the gate.
  - **D L0:** The gate only runs when the COPILOT_AUTO_MODE flag is on, which defaults to False without LaunchDarkly. — [autogpt_platform/backend/backend/copilot/gate/__init__.py:101](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/__init__.py#L101) (verified)
    - *To reach the next level:* Gate is opt-in.
  - **B L0:** Same environment: irreversible external actions with no undo. — [autogpt_platform/backend/backend/blocks/email_block.py:101](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/email_block.py#L101) (verified)
    - *To reach the next level:* No rollback or quantity bounds.
- **Cap:** C2-POWERBYPASS — The generic HTTP request block, the single most general outward action path, skips the review in the default configuration.

### C3 Tool & action scoping — 0.40 (high)

The shared HTTP client is a real SSRF defence: it resolves every DNS answer, blocks private and metadata ranges, pins the connection to the checked IP and re-validates every redirect. Workspace file paths are checked with realpath containment, and the SQL block defaults to read-only. But the default AutoPilot tool set also includes a raw shell, an arbitrary-public-URL HTTP block and a browser whose URL validation does not cover every path, and every tool is on by default.

- **S L2:** Strong URL and path validators exist, but general-purpose tools (raw shell string, any public URL, browser) remain and SQL read-only is keyword-based. — [autogpt_platform/backend/backend/util/request.py:276-283](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/request.py#L276-L283); [autogpt_platform/backend/backend/util/request.py:565](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/request.py#L565); [autogpt_platform/backend/backend/util/request.py:616](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/request.py#L616); [autogpt_platform/backend/backend/copilot/tools/workdir.py:211-213](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/workdir.py#L211-L213); [autogpt_platform/backend/backend/blocks/sql_query_block.py:297-298](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/sql_query_block.py#L297-L298) (verified)
  - *To reach the next level:* General tools are not replaced by narrow ones, and browser URL validation does not cover every path.
- **C L2:** Most built-in network tools use the shared validated client; the shell (when on E2B) is outside it, and browser coverage is incomplete. — [autogpt_platform/backend/backend/copilot/tools/web_fetch.py:208-209](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/web_fetch.py#L208-L209); [autogpt_platform/backend/backend/copilot/tools/bash_exec.py:151](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/bash_exec.py#L151) (verified)
  - *To reach the next level:* E2B shell egress is not covered by the shared validation layer, and browser coverage is incomplete.
- **D L1:** All AutoPilot tools, including shell, browser and every block, are offered by default; tool allow/deny lists exist only for the AutoPilot block inside graphs. — [autogpt_platform/backend/backend/copilot/permissions.py:40-42](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/permissions.py#L40-L42) (verified)
  - *To reach the next level:* No read-only default tool set.
- **B L1:** A misused tool can reach any public host and any service the user connected. — [autogpt_platform/backend/backend/blocks/http.py:71](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/http.py#L71); [autogpt_platform/backend/backend/executor/manager.py:403-416](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/executor/manager.py#L403-L416) (verified)
  - *To reach the next level:* Tools are not quantity-bounded or scoped to one project.
- **Cap:** none

### C4 Code-execution isolation — 0.60 (high)

On a default self-hosted install the AutoPilot shell runs inside bubblewrap: a cleared environment, no network, a read-only system filesystem, a per-session writable workspace, process and memory limits and a 120-second cap. If bubblewrap is missing the tool refuses rather than running on the host. There is no seccomp filter, and the agent-browser's Chromium runs directly in the root backend container with that container's full environment. The E2B cloud sandbox is stronger isolation but is only used when an E2B key is configured, and then receives the user's integration tokens and unrestricted egress.

- **default configuration** (default; raw 0.60 → 0.60) ← counted
  - **S L2:** bubblewrap with a user namespace, cleared env, network unshared and a whitelist filesystem; no seccomp profile or PID namespace. — [autogpt_platform/backend/backend/copilot/tools/sandbox.py:140](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/sandbox.py#L140); [autogpt_platform/backend/backend/copilot/tools/sandbox.py:142](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/sandbox.py#L142); [autogpt_platform/backend/backend/copilot/tools/sandbox.py:188](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/sandbox.py#L188) (verified)
    - *To reach the next level:* No seccomp filter, so the profile does not reach the hardened-sandbox anchor.
  - **C L2:** Shell runs sandboxed and code blocks run only on E2B, but agent-browser launches Chromium as a host subprocess in the backend container. — [autogpt_platform/backend/backend/copilot/tools/bash_exec.py:161](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/bash_exec.py#L161); [autogpt_platform/backend/backend/copilot/tools/agent_browser.py:88](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/agent_browser.py#L88); searched `rg -n '^USER'` in `autogpt_platform/backend/Dockerfile` → 0 hits (No USER directive: the backend container (and the agent-browser Chromium it launches) runs as root.) (verified)
    - *To reach the next level:* The browser path would need to run inside the sandbox too.
  - **D L3:** The sandbox is on by default (bubblewrap is installed in the image) and fails closed when absent; there is no setting that runs the shell on the host. — [autogpt_platform/backend/Dockerfile:120](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/Dockerfile#L120); [autogpt_platform/backend/backend/copilot/tools/bash_exec.py:161](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/bash_exec.py#L161) (verified)
    - *To reach the next level:* The cap of one level above strength prevents a higher rating.
  - **B L3:** Inside bubblewrap: only the session workspace is writable, no secrets in env, no network, ulimits on processes/memory/file size and a 120 s timeout. — [autogpt_platform/backend/backend/copilot/tools/sandbox.py:142](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/sandbox.py#L142); [autogpt_platform/backend/backend/copilot/tools/sandbox.py:188](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/sandbox.py#L188); [autogpt_platform/backend/backend/copilot/tools/sandbox.py:109](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/sandbox.py#L109); [autogpt_platform/backend/backend/copilot/tools/sandbox.py:227](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/sandbox.py#L227) (verified)
    - *To reach the next level:* Workspace persists on the host between runs, so the sandbox is not ephemeral.
- **opt-in E2B cloud sandbox** (alt; raw 0.45, cap G1 → 0.45)
  - **S L4:** A remote ephemeral E2B sandbox replaces local bubblewrap when a key is set. — [autogpt_platform/backend/backend/copilot/tools/bash_exec.py:151](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/bash_exec.py#L151); [autogpt_platform/backend/backend/copilot/config.py:981](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/config.py#L981) (verified)
  - **C L2:** Shell and file tools route to E2B, but the browser still runs in the backend container. — [autogpt_platform/backend/backend/copilot/tools/bash_exec.py:151](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/bash_exec.py#L151); [autogpt_platform/backend/backend/copilot/tools/agent_browser.py:88](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/agent_browser.py#L88) (verified)
    - *To reach the next level:* Browser path not sandboxed.
  - **D L0:** Only active when an E2B key is configured; the shipped .env leaves it blank. — [autogpt_platform/backend/backend/copilot/config.py:981](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/config.py#L981); [autogpt_platform/backend/.env.default:299](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/.env.default#L299) (verified)
    - *To reach the next level:* Opt-in.
  - **B L0:** The sandbox env receives every connected provider token and egress is direct unless the swap proxy address is set. — [autogpt_platform/backend/backend/copilot/tools/bash_exec.py:215](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/bash_exec.py#L215); [autogpt_platform/backend/.env.default:303](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/.env.default#L303) (verified)
    - *To reach the next level:* Credentials in sandbox env with unrestricted network.
- **Cap:** none

### C5 Untrusted input blast radius — 0.23 (high)

Agents read web pages, search results, browser content, MCP tool output, webhooks and messages from other agents, and nothing in the default configuration limits what they can do afterwards. A hijacked AutoPilot can send data out through web fetch or the HTTP block and take irreversible actions through unflagged blocks, and builder graphs triggered by webhooks run unattended. The platform has a content judge that holds suspicious reads for review, but it is an LLM classifier and only runs when the opt-in approval-mode flag is on.

- **default configuration** (default; raw 0.00, cap C5-WORSTCASE → 0.00)
  - **S L0:** No default structural limit on what a session can do after reading untrusted content; fetched content carries no provenance. — [autogpt_platform/backend/backend/copilot/tools/web_fetch.py:208-209](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/web_fetch.py#L208-L209); searched `rg -n -i 'untrusted'` in `autogpt_platform/backend/backend/copilot/tools/web_fetch.py` → 0 hits (Fetched pages are returned to the model with no provenance marking or wrapper.) (verified)
    - *To reach the next level:* Egress and state-changing tools would need to be disabled or approval-gated once untrusted content is read.
  - **C L0:** Untrusted tool results enter context with the same standing as user instructions by default. — [autogpt_platform/backend/backend/copilot/gate/reads.py:219-222](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/reads.py#L219-L222) (verified)
    - *To reach the next level:* Untrusted sources are not distinguished.
  - **D L0:** The only control (the read judge) is off by default because the gate flag is off. — [autogpt_platform/backend/backend/copilot/gate/__init__.py:101](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/__init__.py#L101); [autogpt_platform/backend/backend/copilot/gate/reads.py:219-222](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/reads.py#L219-L222) (verified)
    - *To reach the next level:* Opt-in.
  - **B L0:** A hijacked session can exfiltrate via web_fetch or the HTTP block and take irreversible actions via unflagged blocks; builder graphs on webhook triggers act with no human. — [autogpt_platform/backend/backend/copilot/tools/web_fetch.py:208-209](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/web_fetch.py#L208-L209); [autogpt_platform/backend/backend/blocks/http.py:71](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/http.py#L71); [autogpt_platform/backend/backend/blocks/generic_webhook/triggers.py:23](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/generic_webhook/triggers.py#L23); [autogpt_platform/backend/backend/api/features/library/db.py:570](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/api/features/library/db.py#L570) (verified)
    - *To reach the next level:* Exfiltration and irreversible actions would need human approval.
- **opt-in content judge on outside reads (copilot-auto-mode flag)** (alt; raw 0.23, cap G1 → 0.23) ← counted
  - **S L1:** An LLM content judge holds reads that appear to carry instructions for user review: detection only. — [autogpt_platform/backend/backend/copilot/gate/reads.py:219-222](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/reads.py#L219-L222) (verified)
    - *To reach the next level:* A structural Rule-of-Two limit rather than a classifier.
  - **C L2:** Covers web, browser, shell output, MCP and sub-agent results, but not automation or bot sessions. — [autogpt_platform/backend/backend/copilot/gate/reads.py:41-59](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/reads.py#L41-L59); [autogpt_platform/backend/backend/copilot/gate/__init__.py:97](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/__init__.py#L97) (verified)
    - *To reach the next level:* Automation and bot-platform sessions are not screened.
  - **D L0:** Runs only when the gate flag is on. — [autogpt_platform/backend/backend/copilot/gate/__init__.py:101](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/__init__.py#L101) (verified)
    - *To reach the next level:* Opt-in.
  - **B L0:** Same worst case when the judge misses. — [autogpt_platform/backend/backend/blocks/http.py:71](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/http.py#L71) (verified)
    - *To reach the next level:* Same as default.
- **Cap:** G1 — Opt-in mechanism: off in the scored default configuration.

### C6 Memory, context & configuration integrity — 0.30 (high)

The Claude Agent SDK is started with no settings sources and a strict MCP config, so no project or home-directory files can add tools, hooks or servers. But AutoPilot's self-written skills are on by default: the model can store a skill with no approval, and later sessions receive the skill index as trusted context and can load and follow it. Storage is per-user. Graph memory (Graphiti) is behind an off-by-default flag.

- **S L1:** The model writes skills freely (a workspace effect that runs in every mode) and they are re-injected as a trusted index in later sessions; settings files cannot change security. — [autogpt_platform/backend/backend/copilot/gate/policy.py:83-86](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/gate/policy.py#L83-L86); [autogpt_platform/backend/backend/copilot/service.py:615-618](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/service.py#L615-L618); [autogpt_platform/backend/backend/copilot/sdk/service.py:5143-5144](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/sdk/service.py#L5143-L5144) (verified)
  - *To reach the next level:* Persisted skills would need provenance-as-data presentation or approval.
- **C L1:** Auto-loaded CLI configuration is controlled (setting_sources=[]); skills and business-understanding memory are not. — [autogpt_platform/backend/backend/copilot/sdk/service.py:5143-5144](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/sdk/service.py#L5143-L5144); [autogpt_platform/backend/backend/copilot/tools/skills.py:2312](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/skills.py#L2312) (verified)
  - *To reach the next level:* Skills and understanding stores are not gated.
- **D L2:** Workspaces and skills are namespaced per user in queries. — [autogpt_platform/backend/backend/data/workspace.py:107](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/data/workspace.py#L107) (verified)
  - *To reach the next level:* The cap of one level above strength prevents a higher rating.
- **B L1:** A poisoned skill persists across the user's sessions and can steer tool use. — [autogpt_platform/backend/backend/copilot/tools/skills.py:2312](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/skills.py#L2312); [autogpt_platform/backend/backend/copilot/service.py:615-618](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/service.py#L615-L618) (verified)
  - *To reach the next level:* Not session-scoped and not human-reviewed before persisting.
- **Cap:** none
- **Notes:** Graphiti long-term memory is off unless the graphiti-memory flag is enabled.

### C7 Third-party extensions — 0.30 (high)

The platform loads no third-party code into its own processes: MCP servers are reached only over HTTPS, and agent-browser is version-pinned in the image. Remote MCP servers receive only their own credential and the arguments sent to them. However the model can point run_capability at any HTTPS MCP server URL on its own, server tool lists are fetched fresh with no pinning, and only write-looking tools on non-catalogue servers ask first.

- **S L1:** Remote MCP servers are user- or model-chosen and unpinned; a curated catalogue exists but open-world URLs are accepted. — [autogpt_platform/backend/backend/copilot/tools/run_capability.py:171-172](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/run_capability.py#L171-L172); [autogpt_platform/backend/backend/copilot/capabilities/mcp_review.py:55-56](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/capabilities/mcp_review.py#L55-L56) (verified)
  - *To reach the next level:* No pinning or change detection for MCP tool definitions.
- **C L1:** Only the catalogue distinction applies; open-world servers are unverified. — [autogpt_platform/backend/backend/copilot/tools/run_capability.py:171-172](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/run_capability.py#L171-L172) (verified)
  - *To reach the next level:* Verification for every extension type.
- **D L0:** The model can reach a new MCP server URL without any consent; only write-looking calls on non-catalogue servers pause. — [autogpt_platform/backend/backend/copilot/tools/run_capability.py:171-172](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/run_capability.py#L171-L172); [autogpt_platform/backend/backend/copilot/capabilities/mcp_review.py:55-56](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/capabilities/mcp_review.py#L55-L56) (verified)
  - *To reach the next level:* Adding an extension would need explicit user consent showing what it is.
- **B L3:** MCP servers run remotely and receive only their own stored credential for that server; no local MCP processes are launched. — [autogpt_platform/backend/backend/blocks/mcp/helpers.py:64-74](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/mcp/helpers.py#L64-L74); searched `rg -n -i 'stdio'` in `autogpt_platform/backend/backend/blocks/mcp/block.py autogpt_platform/backend/backend/blocks/mcp/client.py autogpt_platform/backend/backend/copilot/tools/run_mcp_tool.py` → 0 hits (MCP is remote (HTTP) only; no local MCP server process is ever launched.); [autogpt_platform/backend/Dockerfile:150](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/Dockerfile#L150) (verified)
  - *To reach the next level:* Network and data access is not limited to what the extension declares.
- **Cap:** none
- **Notes:** D is L0 because consent is missing, not because a control is opt-in: the catalogue/write-review mechanism is on by default, but the model can reach new MCP servers unprompted. G1 therefore does not apply.

### C8 Secrets & sensitive-data protection — 0.50 (high)

Stored integration credentials are Fernet-encrypted with a per-install key, the backend refuses to boot on a missing or previously published key, OAuth tokens are SecretStr, Sentry events are scrubbed and telemetry is off unless configured. Credentials are injected at block execution and never sent to the model. But the shipped deployment defaults do not lock down credential handling. Subprocesses such as agent-browser inherit the backend's full environment.

- **S L2:** Encrypted-at-rest credentials and SecretStr types, with Sentry scrubbing; but the shipped deployment defaults do not lock down credential handling. — [autogpt_platform/backend/backend/util/encryption.py:26](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/encryption.py#L26); [autogpt_platform/backend/backend/util/secrets_guard.py:49](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/secrets_guard.py#L49); [autogpt_platform/backend/backend/data/model.py:345](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/data/model.py#L345); [autogpt_platform/backend/backend/util/metrics.py:264](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/metrics.py#L264);  (verified)
  - *To reach the next level:* Hardened deployment defaults, and redaction on all major paths including subprocess envs.
- **C L2:** Logs/telemetry are scrubbed and model-bound messages never carry stored credentials; agent-browser and the SDK CLI subprocess inherit the full environment. — [autogpt_platform/backend/backend/util/metrics.py:264](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/metrics.py#L264); [autogpt_platform/backend/backend/copilot/tools/agent_browser.py:88](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/agent_browser.py#L88); [autogpt_platform/backend/backend/copilot/tools/bash_exec.py:55](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/tools/bash_exec.py#L55) (verified)
  - *To reach the next level:* Subprocess environments are not scrubbed.
- **D L3:** Sentry, PostHog and LaunchDarkly are blank by default; scrubbing is always applied when Sentry is configured. — [autogpt_platform/backend/.env.default:268](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/.env.default#L268); [autogpt_platform/backend/.env.default:367](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/.env.default#L367); [autogpt_platform/backend/.env.default:56](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/.env.default#L56) (verified)
  - *To reach the next level:* Stored transcripts are not encrypted or minimised.
- **B L1:** Leaked user OAuth tokens and API keys are long-lived and moderately scoped. — [autogpt_platform/backend/backend/api/features/integrations/router.py:154](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/api/features/integrations/router.py#L154) (verified)
  - *To reach the next level:* Credentials are not short-lived or per-task.
- **Cap:** none

### C9 Audit & traceability — 0.62 (high)

Every chat message, including tool calls and results, is stored in Postgres, graph runs record who or what started them (schedule, webhook, AutoPilot chat) and their parent run, and every human review keeps its status and decision time. This gives a usable trail with actor attribution. The records are ordinary database rows that cascade-delete with the session or user, with no tamper evidence, and tool calls made inside Claude SDK sub-agents only reach application logs.

- **S L3:** Structured records with requesting user, trigger source/ref, parent execution and review decisions. — [autogpt_platform/backend/schema.prisma:733](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/schema.prisma#L733); [autogpt_platform/backend/schema.prisma:1570](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/schema.prisma#L1570); [autogpt_platform/backend/schema.prisma:1600](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/schema.prisma#L1600); [autogpt_platform/backend/schema.prisma:1722-1728](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/schema.prisma#L1722-L1728) (verified)
  - *To reach the next level:* No tamper-evident or off-host storage.
- **C L2:** All built-in tool calls are in the transcript; SDK sub-agent lifecycle is only logged. — [autogpt_platform/backend/schema.prisma:733](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/schema.prisma#L733); [autogpt_platform/backend/backend/copilot/sdk/security_hooks.py:442-456](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/sdk/security_hooks.py#L442-L456) (verified)
  - *To reach the next level:* Sub-agent inner tool calls and credential use are not recorded in the audit store.
- **D L3:** Written by the backend; no model tool deletes transcripts. — [autogpt_platform/backend/schema.prisma:725](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/schema.prisma#L725) (verified)
  - *To reach the next level:* Disabling or deleting is not itself logged; rows cascade-delete.
- **B L2:** Node executions are marked RUNNING in the DB before the block runs. — [autogpt_platform/backend/backend/executor/manager.py:926-929](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/executor/manager.py#L926-L929) (verified)
  - *To reach the next level:* No guarantee every action has a durable replayable record before it executes.
- **Cap:** none

### C10 Limits & kill switch — 0.45 (high)

AutoPilot turns are capped at 100 tool rounds and $10 of model spend, users have daily and weekly cost limits by default, the shell is capped at 120 s, and graph blocks time out after 30 minutes. But graphs have no spend ceiling on a self-hosted install (credits are off), sub-agent limits only gate launches, stopping a graph is cooperative, and schedules keep running after a chat stops.

- **S L2:** Iteration cap plus per-query budget and per-tool timeouts are enforced in code; halt is cooperative. — [autogpt_platform/backend/backend/copilot/config.py:469-470](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/config.py#L469-L470); [autogpt_platform/backend/backend/copilot/config.py:482-483](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/config.py#L482-L483); [autogpt_platform/backend/backend/copilot/sdk/service.py:5153](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/sdk/service.py#L5153); [autogpt_platform/backend/backend/blocks/_base.py:582](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/blocks/_base.py#L582); [autogpt_platform/backend/backend/executor/manager.py:1773-1774](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/executor/manager.py#L1773-L1774) (verified)
  - *To reach the next level:* No wall-clock cap per run and no rate limits on side-effecting tools.
- **C L2:** Top-level AutoPilot loop and tool timeouts are bounded; sub-agent concurrency only gates launches and graph runs started from chat have no spend cap. — [autogpt_platform/backend/backend/copilot/sdk/security_hooks.py:226-234](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/sdk/security_hooks.py#L226-L234); [autogpt_platform/backend/backend/util/settings.py:241-242](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/settings.py#L241-L242) (verified)
  - *To reach the next level:* Sub-agents and spawned graph runs do not share the parent's budget.
- **D L2:** Sensible AutoPilot defaults, operator-configurable; graph spend unlimited by default. — [autogpt_platform/backend/backend/copilot/config.py:412-413](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/copilot/config.py#L412-L413); [autogpt_platform/backend/backend/util/settings.py:241-242](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/settings.py#L241-L242); [autogpt_platform/backend/backend/util/settings.py:299-300](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/settings.py#L299-L300) (verified)
  - *To reach the next level:* Graph runs need a default spend ceiling the model cannot raise.
- **B L1:** Graph runs can spend without limit and scheduled runs continue after a chat stops. — [autogpt_platform/backend/backend/util/settings.py:241-242](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/settings.py#L241-L242); [autogpt_platform/backend/backend/util/settings.py:299-300](https://github.com/Significant-Gravitas/AutoGPT/blob/f8b0e0a87c38b67f6e6cb21f3ee03ff5584c3588/autogpt_platform/backend/backend/util/settings.py#L299-L300) (verified)
  - *To reach the next level:* Tight per-run cost ceilings and stop that cancels scheduled work.
- **Cap:** none

## Rule-of-Two check
[A] untrusted input: web_fetch, browser, MCP results, webhook triggers (autogpt_platform/backend/backend/copilot/tools/web_fetch.py:208) · [B] sensitive data/systems: user's connected OAuth/API credentials injected into blocks (autogpt_platform/backend/backend/executor/manager.py:403) · [C] state change / egress: generic HTTP request block and irreversible blocks (autogpt_platform/backend/backend/blocks/http.py:71) · Same default session? Yes

## Highest-impact improvements
1. Default sensitive_action_safe_mode to True for graphs created in the builder, so scheduled and webhook-triggered runs pause before irreversible blocks. — C2 D L1→L2, +0.050 before caps (Playbook 5)
2. Flag the generic HTTP request blocks (non-GET or any external host) and catalogue MCP write tools for the default review. — C2 C L0→L1, +0.075 before caps (Playbook 5)
3. Ship the copilot-auto-mode gate and read judge on by default, with Ask First for shell and external effects after untrusted reads. — C5 D L0→L2, +0.100 before caps (Playbook 1)
4. Require user approval before a model-written skill is persisted and present skills as untrusted data with origin. — C6 S L1→L2, +0.075 before caps (Playbook 2)
5. Harden the shipped deployment defaults; scrub env for agent-browser. — C8 S L2→L3, +0.075 before caps (Playbook 4)

## Re-audit log
- No changes.

## Limitations
- Static source review of the pinned commit only; nothing was executed, installed, or probed.
- Scored autogpt_platform/ only; the MIT-licensed classic/ AutoGPT agent in the same repo was not reviewed.
- Feature flags were evaluated as on a self-hosted install with no LaunchDarkly key (default values); the managed cloud runs with different flag values.
- Not every one of the ~590 blocks was reviewed individually; block-level findings rely on the shared executor, review and request paths and spot checks.
- Chromium sandbox flags used by agent-browser and the persistence timing of AutoPilot chat messages were not verified.
- No reviewer-steering text was found in the repo's AGENTS.md/CLAUDE.md files.
