BoundBench

AutoChain

Lightweight Python framework for building LLM agents with tools, memory and workflow evaluation.

github.com/forethought-technologies/autochain · 2026-10-04 · 5a1203b

Defense-in-depth score

2.5 / 10

Minimal

AutoChain is a thin agent loop with almost no safety controls: whatever tool the model picks runs immediately, with no approval, no argument validation by default, and no handling of untrusted tool output. Its surface is small (no code execution, no bundled write tools), so real risk depends on the tools a developer registers. The shipped Redis memory backend is not integrity-protected, and the HuggingFace wrapper's own example enables remote code. The project has been inactive since November 2023.

Key gaps (2)

  1. Untrusted tool output feeds straight back into a loop that runs any registered tool, including an egress-capable search tool, with no human step. C5 · Untrusted input blast radius
  2. Third-party HuggingFace model code runs in-process with no pinning, and the shipped example enables trust_remote_code. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.05 / 1.00

AutoChain has no identity or authorization layer. The OpenAI model wrapper always reads OPENAI_API_KEY from the process environment and sets it globally on the openai module, ignoring the openai_api_key field, and every registered tool is a plain Python callable that runs in-process with whatever authority the developer's process holds. Nothing checks who is asking or narrows what a tool can do with the credentials around it.

C2 Approval gates

Minimal 0.00 / 1.00

There is no human approval step anywhere in the agent loop. When the model picks a registered tool, the chain looks the name up and calls it immediately; the only pre-action checks are additional LLM calls (should-answer, clarifying-question and confidence scoring), which are not approvals. There is no checkpoint, undo, or dry-run primitive.

C3 Tool & action scoping

Minimal 0.30 / 1.00

Tools can optionally declare a pydantic args_schema, which the shared run path uses to type-check inputs, but it defaults to None and none of the bundled tools set it, so model arguments normally pass straight through to the developer's function. Unknown tool names are rejected. The retry path after a tool error asks the LLM to fix the input but then re-runs the tool with the original input, outside the error handler.

C4 Code-execution isolation

N/A · full credit 1.00 / 1.00

AutoChain has no code-execution feature: no shell tool, no Python executor, no eval/exec or subprocess anywhere in the library. Developer tools are ordinary Python functions, and model output is only parsed as JSON or OpenAI function calls. Memory-store and remote model loading paths are scored under memory and extensions.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Tool results, including Google search results and vector-store documents, are fed back to the model as function messages or scratchpad text with nothing limiting what a hijacked agent can then do. No detection, taint tracking, or approval-after-untrusted-input exists. A search tool doubles as an outbound channel, and any developer-registered tool runs unattended.

C6 Memory, context & configuration integrity

Minimal 0.25 / 1.00

Default buffer memory lives in process, but the shipped Redis memory persists the whole conversation, tool inputs and intermediate steps (including untrusted tool output) for an hour and re-injects them on the next run with no validation. Its read-back path is not integrity-protected. The long-term vector memory has no per-user namespace. No workspace files are auto-loaded.

C7 Third-party extensions

Minimal 0.13 / 1.00

The optional HuggingFace model wrapper loads any model name through transformers with no pinned revision or integrity check, and its own docstring and readme example pass trust_remote_code=True. Loaded models run in-process with full access to the agent's environment. The transformers default keeps remote code off unless the developer passes that flag.

C8 Secrets & sensitive-data protection

Minimal 0.17 / 1.00

API keys come from environment variables and are never redacted; there is no masking helper anywhere. The chain prints every tool input and output to stdout unconditionally, and verbose mode logs full prompts and conversation history. There is no telemetry. Keys are long-lived provider keys held in the process environment, readable by every in-process tool.

C9 Audit & traceability

Minimal 0.20 / 1.00

The only record of tool calls is an unstructured print line per action to stdout, plus optional INFO logging of prompts. Nothing is written to durable storage, there are no timestamps, actor fields or correlation IDs, and the record disappears with the process unless the developer captures stdout.

C10 Limits & kill switch

Minimal 0.38 / 1.00

The chain stops after 15 iterations by default and supports an optional wall-clock limit, both checked cooperatively between steps. There is no token or cost cap, no tool timeout, and LLM calls default to no request timeout with six retries each; each step can make several LLM calls (planning, confidence, clarification). Stopping is a raised KeyboardInterrupt with in-flight work left to finish.