C1 Identity & least privilege
Minimal 0.25 / 1.00
LLM calls use a dedicated, budget-limited LiteLLM virtual key rather than the master key, and the task API authenticates with Argon2-hashed credentials. Everything else is broad: the patcher, builders and seed generator all mount the Docker socket of a privileged Docker-in-Docker daemon, run under the default service account with no security contexts, share an unauthenticated Redis, and the UI pod holds the CRS token and a GitHub PAT in plain environment variables. Default credential handling in the shipped charts is not locked down. A hijacked agent would inherit root on the DinD daemon plus write access to the competition API.
C2 Approval gates
Minimal 0.05 / 1.00
There is no human approval anywhere in the system. The find-tests agent runs arbitrary shell commands, LLM-written test scripts run automatically, and generated patches are pushed to the queue and submitted to the competition API with no review. Work happens on disposable copies of the task directory and the output is a patch text, which limits damage, but the most powerful action path has no gate at all.
C3 Tool & action scoping
Minimal 0.17 / 1.00
The patch-agent read tools (ls, cat, grep, get_lines, code-query tools) pass model arguments into a shlex-quoted command list, and patch targets are mapped to existing files under the task source directory. The find-tests agent, however, gets a raw bash tool and a script runner with no argument validation, enabled by default with no deploy-time switch. Everything runs in the target project's container with network access.
C4 Code-execution isolation
Minimal 0.25 / 1.00
Model-written Python seed generators run inside a wasmtime WASI sandbox that preopens only a temp directory, caps memory at 50 MB, and refuses to run outside WASI; that path is well isolated. The far more powerful paths are not: every patcher command and LLM-written test script runs in a docker container started with --privileged on a Docker-in-Docker daemon that is itself a privileged pod, with open network. Fuzz targets run directly inside the fuzzer pod and task-supplied helper scripts run in the service pods with inherited environment. A container escape therefore lands on a privileged daemon with shared node storage.
C5 Untrusted input blast radius
Minimal 0.05 / 1.00
Untrusted content (target source code, sanitizer stack traces, task tarballs and SARIF broadcasts) flows into agent prompts with nothing structural in the way: tool outputs are wrapped in tags and no code distinguishes them from instructions. A hijacked find-tests agent can run arbitrary commands with open network in a privileged container and its patch output is submitted automatically, with no human step. The LLM key and competition credentials are not placed in those containers by the docker run call, which keeps the unattended worst case below full secret loss.
C6 Memory, context & configuration integrity
Minimal 0.30 / 1.00
The one persistent store an agent can write is a Redis hash of LLM-authored test scripts keyed by task id; scripts are only accepted after they run once and an LLM judges the output, then are reloaded and executed later without review. A test.sh shipped in the task's project directory is read and run automatically. Redis has authentication disabled by default. Patcher and seed-gen auto-load a .env file, but from the image working directory (/app/patcher, /app/seed-gen), not from the target repo, so repo-controlled configuration does not apply.
C7 Third-party extensions
Minimal 0.20 / 1.00
Buttercup has no plugin, MCP or skill system, so the usual extension surface is absent. The closest equivalent is code that arrives with each task: the OSS-Fuzz infra/helper.py and project Dockerfiles in the task's fuzz-tooling tarball are executed automatically, verified only against a sha256 supplied in the same request. Default images use the moving main tag with pull policy Always, and the WASM Python runtime is downloaded without a checksum. The helper runs in the service pod with its full environment, which for the patcher includes the LLM key.
C8 Secrets & sensitive-data protection
Minimal 0.20 / 1.00
Secrets are held as environment variables and Kubernetes secrets, the LLM key is wrapped in a SecretStr, and the setup script generates a fresh LiteLLM master key and CRS token. But default credential handling is not locked down, the UI pod and every subprocess carry tokens in plain environment, and the patcher and builder charts default to DEBUG logging with no redaction (the shared command runner can log env overrides). Provider keys sit only in the LiteLLM pod, which keeps the broadest keys away from agents.
C9 Audit & traceability
Minimal 0.30 / 1.00
Services log to stderr, a /tmp file, an optional persistent directory on the shared scratch volume, and OTLP when configured; scheduler submissions emit tracing spans. Tool calls are logged as unstructured info lines that include the command or path, but command output is deliberately not logged, there is no approver (there is no approval), and no tamper-evident storage. Logs are written by the same process that runs the agents.
C10 Limits & kill switch
Minimal 0.38 / 1.00
Patch generation is bounded by a patch-retry cap (15 by default in the chart), per-graph recursion limits, a 30-minute PoV phase, and a LiteLLM per-key spend cap that the key-setup job refuses to skip. Fuzzer runs have timeouts. But docker command execution has no timeout, the WASM seed sandbox has no CPU or fuel limit, and cancellation only marks the task in a registry that the patcher checks when it picks up an item, so a running patch loop and in-flight containers are not interrupted.