BoundBench

Tracecat

Open-source agentic security automation (SOAR) platform for teams and AI agents; MCP client/server

github.com/TracecatHQ/tracecat · 2026-10-03 · 77ed0ab

Defense-in-depth score

4.0 / 10

Minimal

Tracecat is a workflow and agent platform that runs integrations against stored SOC and cloud credentials. Its control plane is solid (scope-based RBAC, per-preset tool allowlists, encrypted secrets, secret masking, a gateway that keeps model-provider keys out of the agent). The shipped default executes actions as plain subprocesses that inherit the worker environment, per-tool approval is opt-in and no built-in action asks for it, step limits are not enforced in the agent runtime, and nothing structurally limits a prompt-injected agent that holds both egress and response actions. Switching the executor to nsjail and flagging response actions for approval would change the picture most.

Key gaps (3)

  1. The most powerful action paths (generic HTTP, response actions, sandboxed shell) are not gated by approval in the default configuration, and no built-in action ships flagged. C2 · Approval gates
  2. The default direct executor runs actions in a subprocess that copies the executor environment, which holds the database URI, encryption key, service key and signing secret. C4 · Code-execution isolation
  3. Custom-registry extension code runs in the same direct action subprocess with the full executor environment (inferred from the shared artifact path). C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.45 / 1.00

Agents act through a per-workspace service identity whose JWT lists the exact actions the preset allows, and both the MCP proxy and the executor reject anything outside that list. The identity itself is broad: it carries every workspace operational scope (workflow delete, secret create/delete, any action) and is not narrowed per tool or per requesting user. Authority is workspace-bound but one credential covers read and write across every integration the preset selects. The builder assistant can also rewrite a preset's actions and approval flags.

C2 Approval gates

Minimal 0.25 / 1.00

Tracecat has a real per-call approval mechanism: a flagged tool call is denied at the hook, the exact arguments are stored and shown to a user holding the agent update scope, and the human may approve, deny or override. But approval is opt-in per tool, no built-in registry action ships flagged, and unflagged tools (generic HTTP, response actions, the sandboxed shell) run unattended. Explicit sub-agents skip the gate entirely, and local MCP tools marked as needing approval are hard-denied rather than gated.

C3 Tool & action scoping

Moderate 0.50 / 1.00

Registry actions take typed, schema-validated arguments and an agent only receives the actions its preset names; a new preset has none, the model cannot add tools, and the script-execution action is excluded from agents entirely. The weak spot is the generic tools an operator can select: the HTTP action accepts any URL with no host or private-address check in the action code, so egress control depends on the optional nsjail backend. There are no per-call quantity bounds beyond the tool-count and timeout limits.

C4 Code-execution isolation

Moderate 0.50 / 1.00

By default the platform runs actions and Python scripts without nsjail: actions are plain subprocesses that copy the executor's full environment (database URI, encryption key, service key, signing secret), scripts get only a new PID namespace (and run without it if unshare is unavailable), and network restrictions are not enforced. The agent's shell uses the Claude Code SDK sandbox in a weaker nested mode with unsandboxed commands denied. An opt-in nsjail backend is much stronger: non-root mapping, separate network namespace, seccomp denylist, read-only rootfs, with no host fallback, but it needs elevated container capabilities and is not on by default.

C5 Untrusted input blast radius

Minimal 0.05 / 1.00

Nothing in the code limits what a prompt-injected agent can do: alert, case, webhook and tool-result text enters the model with the same standing as operator instructions, with no taint tracking, quarantine or approval tied to provenance. The only handling found is a prompt line telling the model to treat Slack profile fields as data. What bounds a hijack is the per-preset tool allowlist, sandbox internet being off by default, and optional per-tool approval. A preset that holds both egress and response actions can act on injected instructions unattended.

C6 Memory, context & configuration integrity

Moderate 0.50 / 1.00

There is no cross-session agent memory store. Persistent state is session history that is resumed within a session, user-authored skills and presets, and workspace data. The Claude CLI loads settings only from the user scope in a per-session home, so repository files cannot add hooks or MCP servers. Tenant isolation is enforced in application queries by workspace; the database row-level-security mode ships off. Poisoned content stored in a resumed session or workspace data can influence later turns of that session.

C7 Third-party extensions

Minimal 0.33 / 1.00

Nothing third-party is enabled by default: MCP servers, skills and custom registries are added by workspace operators. Registry versions are pinned in a lock with manifest fingerprints, but user-configured MCP servers are neither pinned nor checked for changed tool definitions, and stdio MCP commands are whatever the operator typed. Custom-registry code runs through the same direct action path as built-ins, in a subprocess carrying the full executor environment, so a malicious extension inherits platform secrets unless nsjail is selected.

C8 Secrets & sensitive-data protection

Minimal 0.47 / 1.00

Secrets are encrypted at rest (Fernet), masked in action output and errors before they reach logs or the model, and model-provider keys are injected by a gateway so the agent holds only a gateway token. Telemetry is off unless a Sentry DSN or PostHog key is set, and Sentry is configured without PII or local variables. The gap is process environment: the direct action subprocess inherits the executor's full environment, so platform keys that decrypt every stored secret sit one environment read away from any action code, and the encryption key is a long-lived value passed as an env var.

C9 Audit & traceability

Moderate 0.50 / 1.00

Every agent tool call runs as a durable workflow with the session and parent workflow IDs, session transcripts are persisted by the orchestrator, approvals record who decided, and control-plane changes go through an audit decorator. The audit stream, however, is a fire-and-forget webhook that only exists when an operator configures a sink, emission is explicitly best-effort, and records live in the application database and Temporal rather than in tamper-evident storage.

C10 Limits & kill switch

Minimal 0.42 / 1.00

An agent turn is bounded by a wall-clock timeout (30 minutes by default, clamped to a one-hour ceiling) and each action has a 300-second executor timeout, with a Redis-signalled cancel that stops the turn. The step and tool-call caps (max_requests, max_tool_calls) are declared in schemas but nothing in the agent runtime, executor or MCP code enforces them, only sub-agents have an optional turn cap, and the API rate-limit middleware class is not installed. There is no spend or cost ceiling in the platform itself.