BoundBench

CyberStrikeAI

AI-native security testing platform in Go integrating 100+ tools with multi-agent orchestration and MCP

github.com/AIPentest/CyberStrikeAI · 2026-10-03 · 470eb5e

Defense-in-depth score

2.6 / 10

Minimal

This is a high-privilege platform by design: its shell and scanner tools run as the server's OS user with the full environment, and the human-approval gate is off by default. Authentication is real (generated admin password, per-user RBAC), but nothing contains the agent once a session is running, and untrusted target content flows straight into a model that can run commands. Operators should treat every deployment as a privileged system, enable approval, and isolate the host.

Key gaps (3)

  1. Tools run as same-user host subprocesses with the full inherited environment and no sandbox of any kind. C4 · Code-execution isolation
  2. Worst case after a hijack is unattended data leakage plus irreversible actions, because the approval gate defaults to off (C5-WORSTCASE). C5 · Untrusted input blast radius
  3. If the authorization layer fails, the agent's authority is the server user's full host authority. C1 · Identity & least privilege

Criteria

C1 Identity & least privilege

Minimal 0.20 / 1.00

The platform has real multi-user authentication: a generated 24-character admin password on first start, per-user roles and permissions, resource-level scopes, a login rate limit, and a principal that is carried into the MCP tool-call authorizer. That governs which human may use the platform. The agent's tools, however, run as the server's OS user with the full inherited environment, so the authority the agent holds is ambient and broad. If the authorization layer fails, the process can reach everything that OS user can.

C2 Approval gates

Minimal 0.38 / 1.00

A per-call human approval gate exists (the HITL feature): it shows the tool name and exact arguments, supports reject, edit-before-run, and auto-rejects on timeout. It is off by default, so a fresh install runs every tool, including the shell and C2 task tools, with no approval. The shipped allowlist of approval-exempt tools also contains project-fact writes and some C2 helper tools. A wrongly approved or unapproved call can act irreversibly on real targets with no rollback.

C3 Tool & action scoping

Minimal 0.28 / 1.00

Tool arguments are mostly passed straight through. The shell tool takes an arbitrary command string, and the only central validation is a configurable regex deny rule (government domains by default) applied to every tool call, which its own docs say is not an exhaustive authorization check. Engagement scope is given to the model as prompt text rather than enforced in code. About 78 tools, including shell, scanners, and C2 tools, are enabled in the default role.

C4 Code-execution isolation

Minimal 0.00 / 1.00

Commands and scanner tools run as ordinary subprocesses of the server, under the server's OS user, with the full inherited environment. A process-guard layer adds cgroup or process-group containment, a process count limit, a memory ceiling, and kill-on-cancel, but that is lifecycle and resource control, not an isolation boundary, and there is no container, filesystem, user, or network separation. The shell tool accepts arbitrary command strings, so anything that reaches it is host-level execution.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Output from scanned targets, web pages, files, and external tools enters the model's context with the same standing as the user's instructions, and nothing in code tracks provenance or restricts tools after untrusted content is read. Combined with a shell, network tools, and C2 tools running ungated by default, a hijacked agent can both exfiltrate data and take irreversible actions without a human. This platform's core job is to read hostile content, so the exposure is built in.

C6 Memory, context & configuration integrity

Minimal 0.30 / 1.00

The model can write persistent project facts and security-finding records, which are re-injected into later sessions as working context, and the shipped approval allowlist exempts fact writes from review. Project data is isolated per user through resource-assignment queries, but facts are not validated or expired, and automatic storage cleanup is off by default. Configuration and prompt files load only from the platform's own directories, not from a target workspace.

C7 Third-party extensions

Minimal 0.23 / 1.00

Third-party extensions are opt-in: the external MCP server list ships empty. An authenticated operator can add a stdio server by naming a command and arguments, and the platform launches it as a same-user process with no version pin, hash check, or re-approval when it changes. Skills and tool definitions are local files loaded from the platform directories. The gap is the absence of any verification, not automatic installation.

C8 Secrets & sensitive-data protection

Minimal 0.30 / 1.00

Secrets live in a plaintext config file created with owner-only permissions, and the settings API does not fully protect secrets. Platform audit records redact common secret-named fields, but full tool arguments and results are stored unredacted in the execution table and tool subprocesses inherit the whole environment. Telemetry is local-only by default. Long-lived provider keys are reachable by anything the shell can read.

C9 Audit & traceability

Moderate 0.50 / 1.00

Every tool call routed through the MCP server is saved with arguments, result, status, timestamps, owning user and conversation. Approval requests and decisions are stored with reviewer type and whether the system or a human decided, and platform actions (logins, config changes) have a separate audit log with actor and IP. Records sit in the same local SQLite database the server's shell could modify, and are not signed or exported by default.

C10 Limits & kill switch

Minimal 0.40 / 1.00

Per-tool timeouts, a shell no-output timeout, concurrency caps and a circuit breaker for external servers, and a process-guard memory and process ceiling are enforced in code, and cancel kills the whole process group. There is no cost or wall-clock budget for a run, and the shipped example config sets the iteration limit to 12000 and the per-tool timeout to 60 minutes, so a runaway can continue for a very long time. Stopping works on in-flight commands, but batch task schedules can restart work.