BoundBench

State of agent safeguards

256 open-source AI agents, scored on their safeguards. The average is 3.1 out of 10.

Key numbers

  • We scored 256 open-source AI agents on how well their own code limits the damage when something goes wrong. The average was 3.1 out of 10.
  • 91% scored below 5. Only two scored 7 or higher.
  • In 64% of agents, a prompt injection could get the agent to leak data and do something irreversible without anyone approving it, on default settings.
  • 98% break Meta's Agents Rule of Two. Within one session they can process untrustworthy inputs, have access to sensitive systems or private data, and change state or communicate externally. The rule allows no more than two of the three.
  • Prompt injection is where agents are weakest. They earned 15% of the available points there, and 39% earned none.
  • 53% already have a safeguard in their code that is turned off by default.
  • Of the agents that run code, 63% do it without a sandbox unless the user sets one up.
  • Popularity doesn't help. The 10 most-starred agents average 3.1, about the same as the rest. OpenClaw, the most-starred agent we scored (392k GitHub stars), got 2.8.
  • Agents from big tech companies did only slightly better: 3.38 on average, against 3.09 for everyone else.

The safeguards agents miss most

CriterionPoints earnedAgents earning none
C5Untrusted input blast radius15%39%
C1Identity & least privilege20%24%
C7Third-party extensions20%13%
C6Memory, context & configuration integrity23%5%
C2Approval gates24%18%
C4Code-execution isolation30%27%
C8Secrets & sensitive-data protection30%4%
C3Tool & action scoping33%5%
C10Limits & kill switch36%2%
C9Audit & traceability38%6%

Critical gaps, by how many agents have them

GapShareAgents
A hijacked agent can leak data and take an irreversible action with no human in the loop64%165
Files in the workspace can turn on tools or loosen safeguards without a trust prompt28%72
The most powerful action (often the shell) skips the approval step21%54
Runs remote code or loads untrusted model files without consent18%46
The agent can widen its own permissions8%21
The model or message content can satisfy the approval step7%19
Forwards a client token to downstream APIs (token passthrough)3%8
The agent is triggered by public events while holding write credentials3%8
Secrets are routinely sent to the model provider2%5

The 10 most-starred agents

AgentGitHub starsScore
OpenClaw391,5882.8
Hermes Agent251,8932.9
DeepSeek Harness (dsh)245,1243.2
OpenCode212,1891.6
n8n206,8193.4
GitHub Copilot Chat (VS Code)193,6303.3
AutoGPT Platform187,6844.2
Dify158,0302.9
Langflow155,5673.0
Open WebUI154,1453.5

Scores from 2026-10-03 to 2026-10-05 on scorecard v0.9. GitHub stars taken on 2026-10-07. The statistics are open data at /data/stats.json.