# Defense-in-Depth Score: OpenSRE (Tracer)

**Repo:** https://github.com/Tracer-Cloud/opensre · **Commit:** `01ba24f1861965e40fa1213956cb86d254ab5373` · **Reviewed:** 2026-10-03
**What it is:** Open-source framework for building AI SRE agents that investigate incidents across 60+ observability/infra tools
**Category:** Infrastructure & Ops
**Scored configuration:** Interactive `opensre` REPL (the mode the README leads with) from a fresh install at the default Auto (High) autonomy level, with integrations as the operator configures them; headless `opensre ask` and the Slack/Telegram gateway are footnoted.
**Agent surface (default):** code execution yes · filesystem write yes · network egress yes · external credentials yes · persistent memory yes · untrusted input yes · third party extensions opt-in · sub agents opt-in · external communication opt-in

## Score: 2.1 / 10.0 (Minimal)

| # | Criterion | S | C | D | B | Raw | Cap | Score | Confidence |
|---|---|---|---|---|---|---|---|---|---|
| C1 | Identity & least privilege | L0 | L0 | L0 | L0 | 0.00 | — | **0.00** | High |
| C2 | Approval gates | L2 | L1 | L0 | L0 | 0.23 | G1 | **0.23** (alt) | High |
| C3 | Tool & action scoping | L1 | L1 | L0 | L0 | 0.15 | — | **0.15** | High |
| C4 | Code-execution isolation | L0 | L0 | L0 | L0 | 0.00 | — | **0.00** | High |
| C5 | Untrusted input blast radius | L1 | L1 | L1 | L0 | 0.20 | C5-WORSTCASE | **0.20** | High |
| C6 | Memory, context & configuration integrity | L0 | L0 | L1 | L1 | 0.10 | — | **0.10** | High |
| C7 | Third-party extensions | L1 | L1 | L1 | L1 | 0.25 | — | **0.25** | Medium |
| C8 | Secrets & sensitive-data protection | L2 | L2 | L0 | L0 | 0.30 | — | **0.30** | High |
| C9 | Audit & traceability | L2 | L2 | L2 | L1 | 0.45 | — | **0.45** | High |
| C10 | Limits & kill switch | L2 | L2 | L1 | L1 | 0.40 | — | **0.40** | High |


As shipped, OpenSRE's interactive agent runs any shell command, GitHub CLI call or chat message the model chooses with no approval prompt, and the code labels this 'alpha mode: allow everything'. Commands execute directly on the host with the operator's full environment, so production cloud, cluster and GitHub credentials are in reach of anything a poisoned log line or alert can talk the model into. Telemetry also sends prompts and tool context to the vendor's analytics by default. Use the headless 'opensre ask' mode, which blocks writes unless explicitly allowed, or set /auto off and run with narrowly scoped, read-only credentials.

## Critical gaps
- The default shell tool runs every command with the operator's full environment (cloud keys, tokens, kubeconfig) and no authorization layer, so a hijacked session holds production-level authority. (ASI03, T3; C1) — [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146); [tools/interactive_shell/shell/execution.py:68-74](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L68-L74)
- In the default REPL (autonomy High) every shell command and every integration write runs without human approval; the code calls this 'Alpha mode: allow everything'. (ASI09, ASI02, T10; C2) — [tools/interactive_shell/shared/execution_policy.py:3](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shared/execution_policy.py#L3); [config/constants/repl_autonomy.py:23](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/config/constants/repl_autonomy.py#L23); [tools/interactive_shell/shell/policy.py:41-46](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/policy.py#L41-L46)
- Model-generated shell commands execute on the host as the operator with the full credential-bearing environment and no sandbox. (ASI05, T11; C4) — [tools/interactive_shell/shell/execution.py:53](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L53); [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146)
- A prompt injection in logs, alerts, issues or chat can make the agent leak credentials and take irreversible production actions unattended. (ASI01, LLM01, T6; C5) — [tools/interactive_shell/shell/policy.py:41-46](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/policy.py#L41-L46); [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146)

## Criterion details

### C1 Identity & least privilege — 0.00 (high)

OpenSRE acts with whatever credentials the operator's machine already holds. The generic AWS tool builds boto3 clients from the default credential chain, the Kubernetes integration loads the user's kubeconfig, and the default shell tool runs every command with the full process environment, so cloud keys, tokens and ~/.aws or ~/.kube files are all within reach. There is no per-request authorization layer; a few integrations (EKS AssumeRole, the scrubbed environment of the Python tool) narrow authority on their own paths only. If the agent is steered wrongly it can do anything the operator's credentials allow, often production-level access.

- **S L0:** Ambient authority: boto3 default chain, the user's kubeconfig, and the full environment handed to shell commands; no scoped identity of the agent's own. — [integrations/aws/aws_sdk_client.py:183](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/aws/aws_sdk_client.py#L183); [integrations/kubernetes/client.py:141](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/kubernetes/client.py#L141); [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146) (verified)
  - *To reach the next level:* No dedicated, role-scoped identity; L1 needs at least a dedicated identity for the agent.
- **C L0:** Each integration constructs its own client from ambient credentials and shell_run passes the whole environment to every command; only the Python tool scrubs its environment and only EKS can use an operator-set AssumeRole. — [tools/interactive_shell/shell/execution.py:68-74](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L68-L74); [infrastructure/safety/sandbox/runner.py:295-304](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/infrastructure/safety/sandbox/runner.py#L295-L304); [integrations/eks/eks_k8s_client.py:28-35](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/eks/eks_k8s_client.py#L28-L35) (verified)
  - *To reach the next level:* No authorization check on any tool path; L1 needs the main tool path to be checked.
- **D L0:** A fresh install runs with the operator's full credentials; narrowing them is manual hardening outside OpenSRE. — [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146); [integrations/aws/aws_sdk_client.py:183](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/aws/aws_sdk_client.py#L183) (verified)
  - *To reach the next level:* Default install inherits full operator privilege; L1 needs a narrower default identity.
- **B L0:** A hijacked session can use the operator's cloud, cluster and GitHub credentials directly through shell_run and github_cli, which for an SRE tool typically means production write access across systems. — [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146); [integrations/github/tools/github_cli/tool.py:126](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/github/tools/github_cli/tool.py#L126); [integrations/kubernetes/client.py:141](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/kubernetes/client.py#L141) (verified)
  - *To reach the next level:* Blast radius is the operator's whole credential set; L1 needs authority limited to bounded write across fewer systems.
- **Cap:** none

### C2 Approval gates — 0.23 (high)

By default nothing asks a human before acting. The REPL ships at autonomy level High, which the code labels 'allow all', and the shell policy allows every command including sudo, pipes and command substitution. Lower levels (/auto med or off) add a per-call prompt, but only for the REPL's own action types such as shell commands; registered integration tools such as github_cli (any gh command, including gh api and merges) and Slack/Telegram send tools never reach that prompt. The headless 'opensre ask' mode is the exception: it blocks every mutating or unknown tool unless explicitly allowed.

- **default configuration** (default; raw 0.00, cap C2-POWERBYPASS → 0.00)
  - **S L0:** Default policy returns allow for every shell command and High asks for no tool type; the module states 'Alpha mode: allow everything'. — [tools/interactive_shell/shared/execution_policy.py:3](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shared/execution_policy.py#L3); [tools/interactive_shell/shell/policy.py:41-46](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/policy.py#L41-L46); [config/constants/repl_autonomy.py:70](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/config/constants/repl_autonomy.py#L70) (verified)
    - *To reach the next level:* No approval at all by default; L1 needs at least a blanket human approval.
  - **C L0:** The most powerful tool, shell_run, is exempt by default, and github_cli is registered with requires_approval=False. — [tools/interactive_shell/shell/policy.py:41-46](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/policy.py#L41-L46); [integrations/github/tools/github_cli/tool.py:113](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/github/tools/github_cli/tool.py#L113) (verified)
    - *To reach the next level:* The shell and gh tools bypass any gate; L1 needs the most powerful tools gated.
  - **D L0:** The session's auto level defaults to High (no prompts); approval exists only if the user opts into a lower level. — [config/constants/repl_autonomy.py:23](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/config/constants/repl_autonomy.py#L23); [surfaces/interactive_shell/session/terminal_session.py:84](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/surfaces/interactive_shell/session/terminal_session.py#L84) (verified)
    - *To reach the next level:* Approval is opt-in; L1 needs it on by default.
  - **B L0:** Unapproved actions include arbitrary shell commands, gh merges/deletes via gh api, and external Slack messages, none reversible by OpenSRE. — [tools/interactive_shell/shell/execution.py:53](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L53); [integrations/github/tools/github_cli/tool.py:91](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/github/tools/github_cli/tool.py#L91); [integrations/slack/tools/slack_send_message_tool/tool.py:49](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/slack/tools/slack_send_message_tool/tool.py#L49) (verified)
    - *To reach the next level:* No checkpoints or undo for irreversible actions; L1 needs at least some actions reversible.
- **opt-in /auto med|low|off confirmation prompts** (alt; raw 0.23, cap G1 → 0.23) ← counted
  - **S L2:** Per-call 'Command to approve' prompt for action tool types, but what the approver sees does not always match the executed content. — [config/constants/repl_autonomy.py:55-75](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/config/constants/repl_autonomy.py#L55-L75) (verified)
    - *To reach the next level:* Approver does not always see the exact executed content; L3 needs the full command/arguments shown with risk tiers.
  - **C L1:** Only REPL action tool types are promoted to ask; gating does not cover registered integration tools, so github_cli and Slack/Telegram send tools still run unprompted. — [integrations/github/tools/github_cli/tool.py:113](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/github/tools/github_cli/tool.py#L113) (verified)
    - *To reach the next level:* Mutating integration tools are not gated; L2 needs every built-in tool gated.
  - **D L0:** Off unless the user runs /auto; /trust also skips prompts. — [config/constants/repl_autonomy.py:23](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/config/constants/repl_autonomy.py#L23); [tools/interactive_shell/shared/execution_policy.py:133-139](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shared/execution_policy.py#L133-L139) (verified)
    - *To reach the next level:* Opt-in; L1 needs on by default.
  - **B L0:** Same irreversible action set as the default. — [integrations/github/tools/github_cli/tool.py:91](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/github/tools/github_cli/tool.py#L91); [integrations/slack/tools/slack_send_message_tool/tool.py:49](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/slack/tools/slack_send_message_tool/tool.py#L49) (verified)
    - *To reach the next level:* No rollback or previews for external actions; L1 needs some reversibility.
- **Cap:** G1 — Opt-in mechanism: off in the scored default configuration.
- **Notes:** Headless `opensre ask` is fail-closed: MUTATING, EXTERNAL and undeclared tools are blocked unless --allowed-tool or --dangerously-bypass-approvals is given (surfaces/cli/ask/approval.py:40-47). Gateway turns gate only tools declaring requires_approval.

### C3 Tool & action scoping — 0.15 (high)

The two most capable default tools are raw passthroughs: shell_run takes any shell string and github_cli takes any gh argument list including gh api. Many integration tools are narrow, fixed read queries, and the generic AWS tool filters operation names with regex allow and block lists, but that is pattern filtering rather than a boundary (get_secret_value still passes). Everything, including shell execution and GitHub writes, is enabled by default.

- **S L1:** Best argument control is regex allow/deny on AWS operation names; shell_run and github_cli accept arbitrary strings. — [integrations/aws/aws_sdk_client.py:14-42](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/aws/aws_sdk_client.py#L14-L42); [tools/interactive_shell/actions/shell.py:61-63](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/actions/shell.py#L61-L63); [integrations/github/tools/github_cli/tool.py:91](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/github/tools/github_cli/tool.py#L91) (verified)
  - *To reach the next level:* Validation is pattern filtering with raw passthrough tools alongside; L2 needs typed schemas with real validation on the powerful tools.
- **C L1:** A few integrations validate or are fixed read queries (AWS op filter, MySQL read_only decorator); the shell and gh tools do not. — [integrations/aws/aws_sdk_client.py:14-42](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/aws/aws_sdk_client.py#L14-L42); [integrations/mysql/__init__.py:173](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/mysql/__init__.py#L173); [tools/interactive_shell/actions/shell.py:61-63](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/actions/shell.py#L61-L63) (verified)
  - *To reach the next level:* Most powerful built-in tools have no validation; L2 needs most built-in tools validating.
- **D L0:** shell_run is available unless a host explicitly zeroes the capability, so write/exec/network tools are on by default. — [core/agent_harness/tools/tool_context.py:82](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/tools/tool_context.py#L82); [tools/interactive_shell/actions/shell.py:121](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/actions/shell.py#L121) (verified)
  - *To reach the next level:* Everything is on by default; L1 needs dangerous tools individually disableable by the operator in normal config.
- **B L0:** shell_run is a general-purpose tool against the whole machine and anything its credentials reach. — [tools/interactive_shell/shell/execution.py:53](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L53); [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146) (verified)
  - *To reach the next level:* General-purpose tools against production; L1 needs at least minor limits on reach.
- **Cap:** none

### C4 Code-execution isolation — 0.00 (high)

Model-written shell commands run directly on the host through /bin/sh with the operator's full environment; there is no container, OS sandbox or separate user. The Python execution tool runs code in a subprocess with a scrubbed environment and monkeypatched open/subprocess/socket, which is filtering rather than isolation, and it is a minor path next to the unrestricted shell. If a command misbehaves it has everything the user's account has, including cloud and cluster credentials.

- **S L0:** Same-user host subprocess via /bin/sh -c; no isolation primitive on the shell path. — [tools/interactive_shell/shell/execution.py:53](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L53); [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146); searched `rg -n -i "docker|firejail|bwrap|landlock|seatbelt|sandbox-exec|nsjail|gvisor|firecracker"` in `tools/interactive_shell` → 2 hits (both hits are 'docker rm'/'docker rmi' strings in the advisory risk-label table (shell/risk.py), not an isolation mechanism) (verified)
  - *To reach the next level:* No isolation; L1 needs at least filtering on the exec path.
- **C L0:** Only the Python tool has a (filtering) restriction; the main exec tool shell_run is not sandboxed. — [infrastructure/safety/sandbox/runner.py:295-304](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/infrastructure/safety/sandbox/runner.py#L295-L304); [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146) (verified)
  - *To reach the next level:* Main exec tool unsandboxed; L1 needs the main exec tool sandboxed.
- **D L0:** No sandbox exists for shell_run in any configuration. — [tools/interactive_shell/shell/policy.py:41-46](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/policy.py#L41-L46); [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146) (verified)
  - *To reach the next level:* No sandbox to enable; L1 needs one on by default.
- **B L0:** Host-equivalent: home directory, ~/.aws, ~/.kube and every credential in the environment are reachable with unrestricted network. — [tools/interactive_shell/shell/execution.py:68-74](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L68-L74); [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146) (verified)
  - *To reach the next level:* Credentials and host files are in reach; L1 needs at least credentials removed from the execution environment.
- **Cap:** none

### C5 Untrusted input blast radius — 0.20 (high)

OpenSRE's job is to read logs, alerts, traces, issues and chat messages that other people and systems wrote, then act with production credentials. Nothing in the default REPL limits what a hijacked session can do: tool results enter the context as ordinary data, there is no taint tracking, and egress (curl via shell, Slack/Telegram sends, gh) and irreversible actions are all unattended. The only spotlighting is a delimiter wrapper for headless 'ask --file' inputs and for task text handed to delegated coding agents. A successful injection can therefore both leak secrets and take destructive actions with no human involved.

- **S L1:** Delimiter/spotlighting only, on two paths (ask --file inputs and coding-agent task text); nothing acts on the labels. — [surfaces/cli/ask/file_input.py:114-126](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/surfaces/cli/ask/file_input.py#L114-L126); [integrations/llm_cli/agent_exec.py:121-136](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/llm_cli/agent_exec.py#L121-L136) (verified)
  - *To reach the next level:* Detection-only; L2 needs dangerous capabilities to require approval after untrusted content is read.
- **C L1:** Integration tool results, shell output and chat messages are not wrapped or distinguished; only the two paths above are. — [surfaces/cli/ask/file_input.py:114-126](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/surfaces/cli/ask/file_input.py#L114-L126); [integrations/llm_cli/agent_exec.py:121-136](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/llm_cli/agent_exec.py#L121-L136) (verified)
  - *To reach the next level:* Most sources unhandled; L2 needs most untrusted sources covered.
- **D L1:** The wrappers are always on where present, but the scored REPL path reads the same content through tools without them. — [integrations/llm_cli/agent_exec.py:121-136](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/llm_cli/agent_exec.py#L121-L136); [tools/interactive_shell/actions/shell.py:61-63](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/actions/shell.py#L61-L63) (verified)
  - *To reach the next level:* Model can route around the wrapper; L2 needs the control on for the default path.
- **B L0:** Hijacked session can exfiltrate via shell network access or Slack sends and take irreversible actions (shell, gh) with no approval. — [tools/interactive_shell/shell/policy.py:41-46](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/policy.py#L41-L46); [integrations/slack/tools/slack_send_message_tool/tool.py:49](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/slack/tools/slack_send_message_tool/tool.py#L49); [integrations/github/tools/github_cli/tool.py:91](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/github/tools/github_cli/tool.py#L91) (verified)
  - *To reach the next level:* Leak plus irreversible action unattended; L1 needs one of those to require a human.
- **Cap:** C5-WORSTCASE — In the default configuration a hijacked session can both exfiltrate secrets and take irreversible actions with no human involved.

### C6 Memory, context & configuration integrity — 0.10 (high)

Long-term memory is on by default: the model can write any fact with memory_remember, an LLM pass extracts 'durable facts' from transcripts after every turn, and the stored memories are injected into every later turn as facts to plan with. The only check on writes is a secret-pattern scan; nothing validates or marks content that came from untrusted sources, so an injection can persist across sessions and steer later tool use. Memories are markdown files under ~/.opensre/memory that the user can inspect and delete. Separately, one integration's configuration loading is not integrity-protected.

- **S L0:** The model can write arbitrary memory content, which is re-injected into every turn as trusted planning facts. — [tools/system/agent_memory/tool.py:94-107](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/system/agent_memory/tool.py#L94-L107); [core/agent_harness/prompts/action/assemble.py:394-396](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/prompts/action/assemble.py#L394-L396) (verified)
  - *To reach the next level:* No gating, provenance or presentation as data; L1 needs at least logged writes with instruction files not elevated.
- **C L0:** Neither the memory tool, the automatic extraction pass, nor the injected index is controlled. — [core/agent_harness/session/memory_extraction.py:1-4](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/session/memory_extraction.py#L1-L4); [core/agent_harness/prompts/action/assemble.py:394-396](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/prompts/action/assemble.py#L394-L396) (verified)
  - *To reach the next level:* No memory path controlled; L1 needs at least one store controlled.
- **D L1:** Memory is per OS user (~/.opensre/memory) and off by default on shared gateway hosts, but it is enabled by default in the CLI and the model writes to it freely. — [core/domain/memory/settings.py:22-24](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/domain/memory/settings.py#L22-L24); [core/domain/memory/settings.py:31-46](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/domain/memory/settings.py#L31-L46) (verified)
  - *To reach the next level:* Isolation is only by storage path with no enforced namespaces on model writes; L2 needs per-user/session namespaces enforced in queries (also limited to S+1).
- **B L1:** Poisoned memory persists across the user's sessions and is injected into every turn where all tools run unapproved. — [core/agent_harness/prompts/action/assemble.py:394-396](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/prompts/action/assemble.py#L394-L396); [tools/interactive_shell/shell/policy.py:41-46](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/policy.py#L41-L46) (verified)
  - *To reach the next level:* Persistent and can trigger tool use; L2 needs it to influence only text or gated actions.
- **Cap:** none
- **Notes:** One integration's configuration loading is not integrity-protected. Not applied as C6-REPOCONFIG because it only fires with that integration configured, which is not the fresh-install default.

### C7 Third-party extensions — 0.25 (medium)

Third-party extensions are MCP servers (Sentry, PostHog, X, GitHub) and local coding-agent CLIs, configured by the operator through settings or environment variables; nothing third-party is enabled out of the box, though one integration's configuration loading is not integrity-protected. MCP stdio servers are launched from operator-supplied commands with no version pin or integrity check (the docs' examples use @latest), and they inherit the agent's entire environment, including every credential. First-party skill releases are auto-updated but are ECDSA-signed and verified against keys compiled into the binary.

- **S L1:** MCP servers run whatever command the operator configured, unpinned; only first-party skill releases are signature-verified. — [integrations/mcp_client.py:89-95](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/mcp_client.py#L89-L95); [core/agent_harness/prompts/skills/snapshot/release.py:167-181](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/prompts/skills/snapshot/release.py#L167-L181) (verified)
  - *To reach the next level:* No pinning or integrity check for MCP servers; L2 needs pinned versions.
- **C L1:** Skill releases are verified; MCP servers and coding-agent CLIs are not. — [core/agent_harness/prompts/skills/snapshot/release.py:167-181](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/prompts/skills/snapshot/release.py#L167-L181); [integrations/mcp_client.py:89-95](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/mcp_client.py#L89-L95) (verified)
  - *To reach the next level:* Only one extension type verified; L2 needs most types.
- **D L1:** MCP servers need an operator-set command (e.g. SENTRY_MCP_COMMAND), but configuration loading on one integration path is not integrity-protected. — [integrations/mcp_client.py:82-87](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/mcp_client.py#L82-L87); [integrations/sentry_mcp/__init__.py:209](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/sentry_mcp/__init__.py#L209) (inferred)
  - *To reach the next level:* Extension configuration is not guaranteed to come only from explicit operator install; L2 needs explicit install only, never workspace-supplied.
- **B L1:** MCP stdio servers are separate processes that receive the full os.environ plus their own settings. — [integrations/mcp_client.py:89-95](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/integrations/mcp_client.py#L89-L95) (verified)
  - *To reach the next level:* Full environment passed; L2 needs a scrubbed environment with only the extension's own configuration.
- **Cap:** none

### C8 Secrets & sensitive-data protection — 0.30 (high)

Stored credentials live in an owner-only (0600) plaintext JSON file rather than an OS keychain, and regex redaction of common token shapes is applied to prompt history, the prompt log, analytics properties and decision traces. However, product telemetry is on by default and ships each turn's prompt, response and up to 48k characters of model context (including tool output) to PostHog, protected only by that regex redaction. Identifier masking and configurable guardrails exist but are off unless enabled. The shell tool inherits every secret in the environment, so the model can read long-lived keys at will.

- **S L2:** Owner-only plaintext credential file plus regex redaction on log and telemetry paths; no keychain. — [config/secrets/local_file.py:114](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/config/secrets/local_file.py#L114); [config/secrets/store.py:10](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/config/secrets/store.py#L10); [config/secret_redaction.py:22-42](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/config/secret_redaction.py#L22-L42) (verified)
  - *To reach the next level:* No keychain or encryption at rest and no redaction before model-bound messages by default; L3 needs both.
- **C L2:** History, prompt log, analytics and traces are redacted; model-bound messages (masking off), WAL transcripts and subprocess environments are not. — [config/prompt_log.py:48](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/config/prompt_log.py#L48); [infrastructure/safety/masking/policy.py:42](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/infrastructure/safety/masking/policy.py#L42); [tools/interactive_shell/shell/execution.py:68-74](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L68-L74) (verified)
  - *To reach the next level:* Model-bound messages and subprocess environments are unprotected; L3 needs all major paths.
- **D L0:** PostHog prompt logging defaults on and sends prompts, responses and model context to the vendor's analytics; opt-out via OPENSRE_NO_TELEMETRY. — [config/prompt_log.py:47](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/config/prompt_log.py#L47); [infrastructure/analytics/prompt_log/recorder.py:409](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/infrastructure/analytics/prompt_log/recorder.py#L409); [infrastructure/analytics/prompt_log/recorder.py:452-453](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/infrastructure/analytics/prompt_log/recorder.py#L452-L453) (verified)
  - *To reach the next level:* Telemetry sends prompt and tool content by default; L1 needs default telemetry to be content-free.
- **B L0:** Long-lived, high-privilege cloud/cluster/GitHub keys are in the environment of every shell subprocess the model starts. — [tools/interactive_shell/shell/execution.py:68-74](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L68-L74); [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146) (verified)
  - *To reach the next level:* Long-lived high-privilege keys reachable by the model; L1 needs them moderately scoped.
- **Cap:** none
- **Notes:** Sentry error reporting is initialized with send_default_pii=False (infrastructure/observability/errors/sentry.py:439). Guardrail rules load only from ~/.opensre/guardrails.yml when present.

### C9 Audit & traceability — 0.45 (high)

Every tool call in the action loop is written to a per-session JSONL file under ~/.opensre/sessions: an intent record with the tool name and arguments is fsynced before the tool runs, and a commit record with the result follows. This gives a replayable local record outside the workspace. It does not attribute actions to a requesting principal or approver, approval decisions go only to analytics, the file is writable by the same user the shell tool runs as, and recording failures are silently swallowed so the action proceeds anyway.

- **S L2:** Structured WAL record of each tool call with arguments, result status and timestamps. — [core/agent_harness/turns/wal_recorder.py:62-70](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/turns/wal_recorder.py#L62-L70); [core/agent_harness/session/persistence/jsonl_store.py:304-322](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/session/persistence/jsonl_store.py#L304-L322) (verified)
  - *To reach the next level:* No actor/approver attribution or cross-agent correlation; L3 needs that.
- **C L2:** All tool calls through the action loop are recorded; approvals/denials and delegated coding-agent actions are not in the record. — [core/agent_harness/turns/wal_recorder.py:62-70](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/turns/wal_recorder.py#L62-L70); [surfaces/interactive_shell/ui/execution_confirm.py:96-115](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/surfaces/interactive_shell/ui/execution_confirm.py#L96-L115) (verified)
  - *To reach the next level:* Approvals and sub-agent actions missing; L3 needs them.
- **D L2:** On by default and stored in ~/.opensre/sessions outside the workspace, but the unsandboxed shell running as the same user can edit or delete it. — [core/agent_harness/session/persistence/paths.py:3](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/session/persistence/paths.py#L3); [tools/interactive_shell/shell/execution.py:137-146](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/shell/execution.py#L137-L146) (verified)
  - *To reach the next level:* Writer is reachable by the model's own tools; L3 needs a component the model cannot control.
- **B L1:** Intents are fsynced per action, but any recording failure is swallowed at debug level and the tool still runs. — [core/agent_harness/turns/wal_recorder.py:82](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/turns/wal_recorder.py#L82); [core/agent_harness/session/persistence/jsonl_store.py:304-322](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/session/persistence/jsonl_store.py#L304-L322) (verified)
  - *To reach the next level:* Logging errors are not surfaced; L2 needs errors surfaced with per-action flush.
- **Cap:** none

### C10 Limits & kill switch — 0.40 (high)

Each turn is capped at 64 tool-calling iterations, shell commands time out after 240 seconds, and cancelling (Ctrl+C/ESC) kills the shell command's whole process group. There is no wall-clock or token/cost ceiling. The model can also extend its own run: session_goal_set accepts an arbitrary max_turns, and the every-10-turns human checkpoint applies only to goals without a turn budget, so a model-chosen large budget skips it.

- **S L2:** Iteration cap plus per-command timeout enforced in code; halt kills the process group. — [core/agent_harness/turns/action_driver.py:107](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/turns/action_driver.py#L107); [tools/interactive_shell/subprocess.py:30](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/subprocess.py#L30); [tools/interactive_shell/subprocess.py:134-137](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/subprocess.py#L134-L137); searched `rg -n -i "max_cost|cost_cap|spend_limit|budget_usd|max_spend"` in `core tools config` → 6 hits (all hits are fleet_monitoring hourly_budget_usd, a reserved/future alarm for monitoring other local agents; none bounds OpenSRE's own spend) (verified)
  - *To reach the next level:* No token/cost or wall-clock cap; L3 needs all three plus rate limits.
- **C L2:** Limits apply to the top-level loop and shell timeouts; goal continuation turns each get a fresh 64-iteration budget. — [core/agent_harness/turns/action_driver.py:107](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/turns/action_driver.py#L107); [tools/interactive_shell/actions/session_goal.py:179-183](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/actions/session_goal.py#L179-L183) (verified)
  - *To reach the next level:* Delegated and continued work does not share one budget; L3 needs sub-agents and spawned work to count against the same budget.
- **D L1:** Defaults are constants, but the model can set a large max_turns on a session goal, which also disables the 10-turn checkpoint. — [tools/interactive_shell/actions/session_goal.py:179-183](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/actions/session_goal.py#L179-L183); [core/agent_harness/session_goal/run_until.py:173-177](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/session_goal/run_until.py#L173-L177) (verified)
  - *To reach the next level:* Model can raise its own multi-turn limit; L2 needs sensible defaults it cannot raise.
- **B L1:** Without a cost ceiling a model-set goal budget can run many 64-iteration turns; stopping does kill in-flight shell commands. — [core/agent_harness/session_goal/run_until.py:173-177](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/core/agent_harness/session_goal/run_until.py#L173-L177); [tools/interactive_shell/subprocess.py:134-137](https://github.com/Tracer-Cloud/opensre/blob/01ba24f1861965e40fa1213956cb86d254ab5373/tools/interactive_shell/subprocess.py#L134-L137) (verified)
  - *To reach the next level:* Ceilings are effectively unbounded; L2 needs moderate ceilings.
- **Cap:** none

## Rule-of-Two check
[A] untrusted input: Logs, alerts, issues and chat read via integration tools and shell output (tools/interactive_shell/actions/shell.py:61) · [B] sensitive data/systems: Operator's cloud/cluster/GitHub credentials in the shell environment (tools/interactive_shell/shell/execution.py:68-74) · [C] state change / egress: shell_run and github_cli run unapproved (tools/interactive_shell/shell/policy.py:41-46, integrations/github/tools/github_cli/tool.py:113) · Same default session? Yes

## Highest-impact improvements
1. Default the REPL to /auto med (ask for shell and mutating tools) and route every MUTATING/EXTERNAL registered tool, including github_cli and chat sends, through the same prompt, as headless ask already does. — C2 D L0→L2, +0.100 before caps (Playbook 5, step 1)
2. Pass shell_run a scrubbed environment (like the Python tool's _sandbox_env) and inject only the credentials a given integration needs. — C8 B L0→L1, +0.050 before caps (Playbook 4)
3. Make PostHog prompt/response/context capture opt-in and keep default telemetry content-free. — C8 D L0→L2, +0.100 before caps
4. Run shell_run inside an OS sandbox (Seatbelt/Landlock or a hardened container) with workspace-only writes and egress off by default. — C4 S L0→L3, +0.225 before caps (Playbook 3, step 1)
5. Mark memories with their source, show them to the model as data, and require user confirmation for model-initiated memory writes. — C6 S L0→L2, +0.150 before caps (Playbook 2)

## Re-audit log
- No changes.

## Limitations
- Static source review of the pinned commit only; nothing was executed, installed, or probed.
- Scored the interactive REPL default; headless `opensre ask` (fail-closed approvals) and the Slack/Telegram gateway (allowlisted senders, requires_approval-only gating) would score differently on C2 and C5.
- The 60+ individual integration tools were sampled (AWS, Kubernetes/EKS, MySQL, GitHub, Slack, MCP) rather than read exhaustively; the hosted OpenSRE Cloud service and webapp are not in this repo and were not examined.
- No reviewer-directed instructions were found in AGENTS.md, CLAUDE.md, README.md or SECURITY.md.
