BoundBench

PR-Agent

Open-source AI PR reviewer (GitHub/GitLab/Bitbucket bot, Action, CLI)

github.com/The-PR-Agent/pr-agent · 2026-10-03 · 090a7e4

Defense-in-depth score

5.3 / 10

Moderate

PR-Agent's model has no tools: deterministic code publishes its structured output to the triggering PR, it runs no code, loads no plugins, and reads its configuration only from the default branch. The dominant risk is who can drive it: authorization of comment-triggered commands and of comment-supplied setting overrides is not locked down. Nothing is human-approved before publishing.

Criteria

C1 Identity & least privilege

Moderate 0.50 / 1.00

In the GitHub Action, PR-Agent acts with the workflow's GITHUB_TOKEN, a repository-scoped token that lives only for the job. The official workflow grants it write access to contents, pull requests, issues and checks, and one token serves every read and write. Apart from one per-user access check for sibling-repository context files, there is no authorization check in code: whoever can trigger the workflow gets the token's full authority.

C2 Approval gates

Minimal 0.10 / 1.00

PR-Agent has no human approval step. When a PR opens, it automatically rewrites the PR description, posts a review and code suggestions, and sets labels. Any comment command runs straight away, including one that commits a CHANGELOG.md update to the PR branch. Merging and approving are not possible: the approval options are disabled and blocked as comment arguments. Most of what it does can be undone, but comments and notifications go out immediately.

C3 Tool & action scoping

Moderate 0.57 / 1.00

The model has no tools. Deterministic code decides each action and publishes the model's structured output to the same PR, so the set of actions is narrow and bounded by design. The weak spot is comment arguments: filtering of comment-supplied setting overrides is not a strict boundary. Every command passes through this one filter.

C4 Code-execution isolation

N/A · full credit 1.00 / 1.00

PR-Agent never runs model-written or PR-supplied code. The model returns text that is parsed into a fixed review format, and the only subprocesses are fixed-argument shallow git clones. Prompt templates are rendered with Jinja's sandboxed environment, and their source comes from host or default-branch settings that comment arguments cannot change. There is no code-execution surface to isolate.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

Untrusted content reaches the model from PR diffs, titles and descriptions, linked issues and comments. Because the model has no tools, a hijacked run can only change the text PR-Agent publishes to that PR; it cannot choose actions or destinations. The bigger gap is authorization of who can drive the bot through comments, which is not locked down.

C6 Memory, context & configuration integrity

Moderate 0.68 / 1.00

PR-Agent loads its settings (.pr_agent.toml) and AGENTS.md context from the repository's default branch, not from the PR, so a pull request cannot rewrite its own reviewer's instructions. Host-only keys such as output sinks, remote config URLs and sibling repositories are stripped from repository settings. The only state it carries between runs is a review-findings marker, read only from comments PR-Agent itself wrote, scoped to that PR and visible in the thread. AGENTS.md is still loaded silently as instruction context.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

PR-Agent loads no third-party code at runtime. It has no plugin system and no MCP client. The optional 'skills' feature reads only text files from host-configured paths, and the only dynamic imports load the built-in git-provider modules.

C8 Secrets & sensitive-data protection

Minimal 0.38 / 1.00

Credentials (the GitHub token and model API key) come from environment variables and never enter the prompt; the model has no tool to read them. The code masks credentials in a few places: clone URLs, the /config listing, and merged settings, which it never logs. There is no general log redaction, and git subprocesses inherit the full environment. Telemetry is off by default, and while the default log level is DEBUG, the console sink drops the prompt payloads. The model API key is long-lived.

C9 Audit & traceability

Minimal 0.38 / 1.00

Each command is logged to stdout with its command name and PR URL, along with every setting a comment changed; GitHub keeps the Action log outside anything the agent can write. The console logs are unstructured, and the code does not record who requested a command. Everything PR-Agent publishes is also visible in the PR's own GitHub history. If logging fails, the action carries on regardless.

C10 Limits & kill switch

Minimal 0.40 / 1.00

Each model call has a 120-second timeout, input is capped at 32k tokens, and chunked tools stop after a few calls. However, these values are not fully protected from comment arguments, there is no spend ceiling, and nothing limits how often commands run. Every new comment starts a fresh run with a fresh budget. Cancelling the GitHub workflow stops the run.