BoundBench

Superagent SDK

Open-source AI safety SDK (TypeScript/Python), CLI and MCP server offering prompt-injection guard, PII redaction, and an LLM-agent repository scanner.

github.com/superagent-ai/superagent · 2026-10-04 · c9e6462

Defense-in-depth score

2.1 / 10

Minimal

Guard and redact are single LLM calls with a well-hardened URL fetcher, but scan is an unattended agent: it installs opencode-ai@latest in a remote Daytona sandbox, points it at an untrusted repository, and gives it the operator's Anthropic and OpenAI API keys with unrestricted network. A poisoned repository can steal those keys, and nothing approves, logs, or time-limits the run. The sandbox keeps the operator's machine safe, but not their keys.

Key gaps (3)

  1. Scan injects the operator's Anthropic/OpenAI API keys into the sandbox where an agent reads an untrusted repository with network access, so a poisoned repo can exfiltrate them. C4 · Code-execution isolation
  2. OpenCode runs with the scanned (untrusted) repository as its working directory, so repo-controlled agent config and instruction files can reconfigure the agent holding the operator's keys. C6 · Memory, context & configuration integrity
  3. Every scan installs and runs opencode-ai@latest unpinned, giving whatever is currently published the operator's LLM keys. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.13 / 1.00

The SDK reads the operator's own API keys from environment variables and, for repository scans, copies the operator's full Anthropic and OpenAI API keys into the remote sandbox where an autonomous agent reads the untrusted repository. Nothing narrows these keys to the task, issues short-lived credentials, or checks authority per request. Forwarding is limited to those two named variables rather than the whole environment, which is the only narrowing present. A hijacked scan agent holds billing-capable provider keys for two vendors.

C2 Approval gates

Minimal 0.05 / 1.00

There is no human approval step anywhere. A scan call (from code, the CLI, or the MCP tool an AI host can invoke) provisions a paid remote sandbox, installs software, and runs an autonomous coding agent with shell and network access, with no confirmation. The only restraint is a prompt telling the scan agent to use read-only tools, which is not a control. The MCP server labels the scan tool read-only and idempotent, so hosts that auto-approve read-only tools will run it without asking.

C3 Tool & action scoping

Minimal 0.38 / 1.00

The guard feature's URL fetcher is well built: it resolves hostnames, rejects private and internal addresses, pins the resolved address, re-validates every redirect, and caps size and time. The scan path is the opposite: the repository URL is only prefix-checked, and scan argument handling is not locked down. The scan agent itself gets OpenCode's general shell, edit and fetch tools. Everything is scoped to a throwaway remote sandbox.

C4 Code-execution isolation

Strong 0.80 / 1.00

All code execution happens in a remote, per-scan Daytona sandbox that is deleted afterwards; the SDK runs nothing on the operator's machine and has no fallback to local execution. That is a strong boundary that the model cannot switch off. What sits inside it is the problem: the operator's Anthropic and OpenAI API keys are injected into the sandbox environment and no network restriction is requested, so code from the scanned repository or a hijacked agent can read and send those keys out.

C5 Untrusted input blast radius

Minimal 0.05 / 1.00

Scan exists to read untrusted repositories, and it hands that content to an autonomous agent with a shell, network access, and the operator's API keys, unattended. The only defence is a prompt telling the agent to stay read-only. Superagent's own guard classifier is not applied to the scanned content, and even as a product it is a detection layer, not a complete boundary. The agent's report is returned to the caller (or an MCP host model) as plain text with no marking that it was derived from attacker-controlled content.

C6 Memory, context & configuration integrity

Minimal 0.20 / 1.00

The project keeps no memory or persistent store of its own. However, the scan agent is started inside the freshly cloned, untrusted repository with no flags isolating it from project configuration, so files in that repository that OpenCode auto-loads (instruction files and project config that can add tools or MCP servers, change permissions, or redirect the model endpoint) can reconfigure the agent that holds the operator's keys. The damage is limited to one scan because each sandbox is deleted afterwards.

C7 Third-party extensions

Minimal 0.05 / 1.00

Every scan runs `npm i -g opencode-ai@latest` inside the sandbox, so whatever version of that third-party agent is newest at that moment is installed and run with the operator's API keys, with no version pin, hash check, or notice to the user. A compromised release would get those keys on the next scan. The sandbox keeps it off the operator's machine.

C8 Secrets & sensitive-data protection

Minimal 0.20 / 1.00

Keys come from environment variables and are not placed in prompts, but the scan path pushes the Anthropic and OpenAI keys into the sandbox where the model-driven agent can read them; the Python SDK's key handling is not locked down either. Usage telemetry is on by default and sends only token counts (with the Superagent key) to superagent.sh, with no opt-out. Error messages embed raw provider responses. There is no redaction of the SDK's own logs or errors.

C9 Audit & traceability

Minimal 0.00 / 1.00

No record of what the scan agent did is kept. The SDK parses OpenCode's event stream and keeps only the final text and token counts, discarding the tool-call events, and the sandbox (with anything it logged) is deleted afterwards. There is no logging module, audit trail, or telemetry of actions.

C10 Limits & kill switch

Minimal 0.20 / 1.00

The guard URL fetcher has firm size and 30-second limits, and MCP inputs are capped at 50,000 characters. The scan, the only agent loop, has no step, time, or cost limit: the sandbox command runs as long as the agent keeps going, spending on the operator's keys. Stopping cleanly deletes the sandbox, but if the process is killed the remote sandbox is left behind.