State of agent safeguards
256 open-source AI agents, scored on their safeguards. The average is 3.1 out of 10.
Key numbers
- We scored 256 open-source AI agents on how well their own code limits the damage when something goes wrong. The average was 3.1 out of 10.
- 91% scored below 5. Only two scored 7 or higher.
- In 64% of agents, a prompt injection could get the agent to leak data and do something irreversible without anyone approving it, on default settings.
- 98% break Meta's Agents Rule of Two. Within one session they can process untrustworthy inputs, have access to sensitive systems or private data, and change state or communicate externally. The rule allows no more than two of the three.
- Prompt injection is where agents are weakest. They earned 15% of the available points there, and 39% earned none.
- 53% already have a safeguard in their code that is turned off by default.
- Of the agents that run code, 63% do it without a sandbox unless the user sets one up.
- Popularity doesn't help. The 10 most-starred agents average 3.1, about the same as the rest. OpenClaw, the most-starred agent we scored (392k GitHub stars), got 2.8.
- Agents from big tech companies did only slightly better: 3.38 on average, against 3.09 for everyone else.
The safeguards agents miss most
| Criterion | Points earned | Agents earning none | |
|---|---|---|---|
| C5 | Untrusted input blast radius | 15% | 39% |
| C1 | Identity & least privilege | 20% | 24% |
| C7 | Third-party extensions | 20% | 13% |
| C6 | Memory, context & configuration integrity | 23% | 5% |
| C2 | Approval gates | 24% | 18% |
| C4 | Code-execution isolation | 30% | 27% |
| C8 | Secrets & sensitive-data protection | 30% | 4% |
| C3 | Tool & action scoping | 33% | 5% |
| C10 | Limits & kill switch | 36% | 2% |
| C9 | Audit & traceability | 38% | 6% |
Critical gaps, by how many agents have them
| Gap | Share | Agents |
|---|---|---|
| A hijacked agent can leak data and take an irreversible action with no human in the loop | 64% | 165 |
| Files in the workspace can turn on tools or loosen safeguards without a trust prompt | 28% | 72 |
| The most powerful action (often the shell) skips the approval step | 21% | 54 |
| Runs remote code or loads untrusted model files without consent | 18% | 46 |
| The agent can widen its own permissions | 8% | 21 |
| The model or message content can satisfy the approval step | 7% | 19 |
| Forwards a client token to downstream APIs (token passthrough) | 3% | 8 |
| The agent is triggered by public events while holding write credentials | 3% | 8 |
| Secrets are routinely sent to the model provider | 2% | 5 |
The 10 most-starred agents
| Agent | GitHub stars | Score |
|---|---|---|
| OpenClaw | 391,588 | 2.8 |
| Hermes Agent | 251,893 | 2.9 |
| DeepSeek Harness (dsh) | 245,124 | 3.2 |
| OpenCode | 212,189 | 1.6 |
| n8n | 206,819 | 3.4 |
| GitHub Copilot Chat (VS Code) | 193,630 | 3.3 |
| AutoGPT Platform | 187,684 | 4.2 |
| Dify | 158,030 | 2.9 |
| Langflow | 155,567 | 3.0 |
| Open WebUI | 154,145 | 3.5 |
Scores from 2026-10-03 to 2026-10-05 on scorecard v0.9. GitHub stars taken on 2026-10-07. The statistics are open data at /data/stats.json.