BoundBench

Agent Governance Toolkit

Microsoft's multi-language toolkit of policy enforcement, identity, audit, sandboxing and SRE primitives for governing AI agent tool calls.

github.com/microsoft/agent-governance-toolkit · 2026-10-04 · c3e8229

Defense-in-depth score

3.5 / 10

Minimal

AGT's govern() wrapper is a real, deterministic, deny-by-default policy gate with careful fail-closed handling, but almost every other safety primitive in the toolkit is opt-in. With default arguments there is no human approval, no sandbox, no untrusted-input handling, no budgets, and an audit log that lives only in memory and omits call arguments. The README's lead policy sets default_action: allow, which turns the gate into a denylist. The main risk is assuming the toolkit's broad feature list is active when only the policy check is.

Key gaps (2)

  1. Governed tools and the core ExecutionSandbox run in the host process; the sandbox is self-described as soft, so an escape reaches the host process's credentials, network and files. C4 · Code-execution isolation
  2. No untrusted-input control is wired into govern(); a hijacked agent can exfiltrate data and take irreversible actions through governed tools unless the developer's own policy forbids it. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.17 / 1.00

govern() puts a deterministic, deny-by-default policy check in front of each wrapped tool call, but it does nothing to narrow the credentials the tool itself uses. The default agent identity is the wildcard "*", so no per-agent or per-user authority is checked unless the developer sets one, and the toolkit's short-lived agent credentials are for agent-to-agent trust, not for the external systems tools reach. Wrapped tools and any subprocesses they start keep the host application's full environment and credentials.

C2 Approval gates

Minimal 0.28 / 1.00

Human approval only happens when a developer writes a policy rule with action require_approval; nothing is gated by default. When it does fire, the safe fallback is to auto-reject, and an optional coordinator binds the approval to a digest of the exact call. But the built-in console and webhook approvers only show the rule, agent and action label, not the actual arguments. The README's lead policy example sets default_action: allow, so everything not explicitly listed runs without a human.

C3 Tool & action scoping

Minimal 0.45 / 1.00

The policy engine is a single central layer that every govern()-wrapped call passes through, and it fails closed: unmatched calls are denied by default, and malformed or unevaluable deny rules count as matches. Its conditions are simple field-versus-literal string and number comparisons, with no resolved-path containment or URL/host checks, and the scalar arguments it sees are whatever the model passed. Tool names never enter the policy context. The README leads with a default-allow policy, which turns the deny-by-default into a denylist.

C4 Code-execution isolation

Moderate 0.50 / 1.00

govern() runs wrapped tools directly in the host process, and the core package's code-execution sandbox is an in-process restricted exec that its own authors call a soft sandbox, not a boundary. The separate agent-sandbox package ships a properly hardened Docker provider: non-root, all capabilities dropped, no-new-privileges, seccomp, read-only root filesystem and no network by default. But it is opt-in, covers only code sent to it, and falls back to a stock image with a warning when the hardened image is missing.

C5 Untrusted input blast radius

Minimal 0.15 / 1.00

Nothing in the default govern() path distinguishes untrusted content: tool results go straight back to the caller, and no taint tracking or Rule-of-Two enforcement exists. The toolkit ships pattern-based prompt-injection and MCP-response scanners, but they are detection-only and must be called separately. Because governed tools can read untrusted content, touch private data and send data out in the same session, a hijacked agent's worst case is limited only by the developer's own policy rules.

C6 Memory, context & configuration integrity

Moderate 0.55 / 1.00

The toolkit does not own an agent memory store, and govern() loads policy only from a path or object the developer passes in, with nothing auto-loaded from the working directory. A MemoryGuard helper can screen memory writes, but it is pattern-based and opt-in. One utility, discover_agents(), reads AGENTS.md and security.md from a repository and maps them to kernel policies, so a developer who points it at an untrusted repo lets that repo define its own security settings.

C7 Third-party extensions

Minimal 0.38 / 1.00

govern() itself loads no third-party code. The plugin marketplace installer requires an Ed25519 signature from a trusted author by default and runs plugins in a subprocess with a minimal environment. But the signature covers the manifest rather than a code digest, the MCP proxy launches whatever server command the user gives it with the full environment, and installed provider packages are auto-loaded in-process through entry points.

C8 Secrets & sensitive-data protection

Minimal 0.30 / 1.00

govern() keeps logs lean: its audit entries record the action label, outcome and matching rule, not the call arguments. But it does no redaction of its own and nothing keeps secrets out of wrapped tools, their subprocesses, or custom approval handlers (which receive the full call context). A credential redactor exists and is applied in the MCP gateway's audit path. OpenTelemetry export is opt-in via bootstrap_otel().

C9 Audit & traceability

Minimal 0.38 / 1.00

Audit is on by default: every governed call writes a policy decision (and every approval or denial) into a Merkle hash-chained log before the tool runs, and a failed audit write stops the call. But the default log lives only in process memory and is lost on exit, and entries record the action label, outcome and matching rule, not the arguments or the tool's result. An HMAC-signed append-only file sink is available if the developer configures a path and key.

C10 Limits & kill switch

Minimal 0.30 / 1.00

govern() has no step, time or cost limits. The only built-in bound on the governed path is an optional per-rule rate limit declared in policy YAML. The toolkit separately ships a token and cost BudgetTracker and a callback-based KillSwitch, but the developer has to wire both into their own loop, and the kill switch relies on the agent's registered callback to actually stop work.