BoundBench

Hermes Agent

Self-improving personal agent with skills, memory, terminal and messaging gateways

github.com/NousResearch/hermes-agent · 2026-10-03 · f4d3e62

Defense-in-depth score

2.9 / 10

Minimal

Hermes Agent runs shell commands and Python directly on the host as your user by default, with every core tool (terminal, code execution, file writes, browser, messaging, cron, sub-agents) enabled. Its approval gate is a pattern denylist whose flagged commands are first judged by an auxiliary LLM, and the execute_code tool skips it entirely in the interactive CLI, so a hijacked model can run arbitrary code, rewrite the agent's own security config, and exfiltrate data without a human. The project's own SECURITY.md says the only real boundary is OS-level isolation (Docker/remote backend or a whole-process sandbox), which is opt-in. Good engineering is visible in env scrubbing, SSRF guards, gateway allowlists and secret redaction, but these are heuristics on an unsandboxed host.

Key gaps (6)

  1. The model can disable the approval gate at runtime: execute_code runs ungated in the CLI and approvals.mode is re-read live from ~/.hermes/config.yaml. C2 · Approval gates
  2. execute_code, arbitrary host Python with subprocess access, skips the approval gate in the default interactive CLI. C2 · Approval gates
  3. Shell and Python execution run directly on the host by default with no isolation. C4 · Code-execution isolation
  4. A hijacked agent holds the user's full host account and every configured service credential. C1 · Identity & least privilege
  5. After reading untrusted content the agent can both exfiltrate secrets and take irreversible actions with no human in the loop. C5 · Untrusted input blast radius
  6. Plugins and skills run in-process with full agent privileges and credentials. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.15 / 1.00

Hermes runs with the full authority of the OS user who launched it. It strips its own provider keys and gateway tokens from the environment of shell, code and MCP child processes, but deliberately leaves the operator's ambient credentials (AWS chain, SSH, gh login) available, and plugins, skills and hooks run inside the agent process with everything it holds. Messaging gateways deny unknown senders by default, but within the allowlist every caller has the same full authority. A hijacked agent therefore holds the user's entire account across every connected service.

C2 Approval gates

Minimal 0.17 / 1.00

Shell commands are checked against a large regex denylist plus the Tirith scanner; anything not flagged runs immediately. Flagged commands go first to an auxiliary-LLM 'smart' approver (the shipped default), which can approve on its own, and only uncertain or denied cases reach the human, who then sees the exact command. The execute_code tool, which runs arbitrary Python on the host, is not gated at all in the interactive CLI, and because the approval mode is read live from ~/.hermes/config.yaml, model-written Python can switch approvals off mid-session. File writes, messaging, MCP tools and Home Assistant calls are not gated, and checkpoints are off by default.

C3 Tool & action scoping

Minimal 0.20 / 1.00

The core tools are general-purpose: a raw shell, arbitrary Python execution, a full browser and unrestricted file writes. Some tools have real validation: web and browser fetches use an SSRF guard that blocks private and cloud-metadata addresses and rechecks redirects, and file writes refuse system paths and the Hermes config. But the shell and Python tools make those checks easy to step around, and every core tool is on by default.

C4 Code-execution isolation

Minimal 0.33 / 1.00

By default every shell command and execute_code script runs directly on the host as the user, with no isolation. A Docker backend is available on request: it drops most capabilities, sets no-new-privileges and PID limits, and runs shell, file and code tools inside the container, but keeps network on, runs as root inside by default, and leaves MCP servers, plugins, hooks and skills on the host. The project itself says OS-level isolation is the only real boundary.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

Only allowlisted principals can instruct the agent through messaging gateways, and webhook-triggered sessions get a read-only tool set. Inside a normal session, results from web search, web extract, browser and MCP tools are wrapped in untrusted-data delimiters, but nothing in code acts on that marking: files, terminal output, email bodies and other tool results are not marked at all, and after reading hostile content the agent keeps full shell, code execution, outbound HTTP and messaging tools. A successful injection can therefore both leak secrets and take irreversible actions with no human involved.

C6 Memory, context & configuration integrity

Minimal 0.25 / 1.00

The agent writes its own long-term memory (MEMORY.md and USER.md, injected into every future system prompt) and its own skills without approval by default; memory writes pass a regex injection scan, while agent-written skills are not scanned by default. Project instruction files such as AGENTS.md, CLAUDE.md and .cursorrules load automatically from the working directory after the same regex scan. Repo-controlled plugins load only behind an explicit environment flag, and no project file can add MCP servers or hooks. Memory is per profile and shared by every session and allowlisted gateway user of that profile.

C7 Third-party extensions

Minimal 0.25 / 1.00

No third-party extension is enabled by default, project-directory plugins need an explicit environment flag, and hub skills are scanned and recorded with a content hash at install. MCP servers launched through npx/uvx get a fail-open malware lookup but are otherwise whatever version the user's command resolves. Plugins and skills run inside the agent process with all of its credentials and tools; MCP stdio servers at least get a scrubbed environment.

C8 Secrets & sensitive-data protection

Minimal 0.40 / 1.00

Secrets live in plaintext ~/.hermes/.env (permissioned 0600), with optional 1Password/Bitwarden sources. A regex redactor runs on log output and on terminal and file content returned to the model, Hermes-managed keys are scrubbed from child environments, and telemetry is off by default. Redaction can be switched off in config, and the keys themselves stay readable by model-run Python, so a hijacked agent can still reach long-lived provider and messaging credentials.

C9 Audit & traceability

Moderate 0.50 / 1.00

Every message, including assistant tool calls and tool results, is stored in a SQLite session database under ~/.hermes with timestamps, and sub-agent sessions link to their parent. The tool-call turn is flushed before tools run, so a crash still leaves a record. Approvals and denials appear only as tool-result text and log lines, there is no actor attribution or tamper evidence, and the database sits where the agent's own shell and Python can modify it.

C10 Limits & kill switch

Minimal 0.40 / 1.00

Per-command terminal timeouts (180 seconds) and an interrupt that kills the foreground process group are on by default, and an iteration cap and wall-clock budget exist but are both unlimited by default. There is no token or spend cap. Sub-agents get their own fresh iteration budget that the model cannot raise, with concurrency and depth limits. The 'hermes pause' emergency stop only blocks new work, and cron jobs and background processes the agent created keep running after a stop.