BoundBench

JARVIS (HuggingGPT)

Microsoft research system in which an LLM plans tasks and dispatches them to Hugging Face models (HuggingGPT), plus EasyTool and TaskBench research code.

github.com/microsoft/JARVIS › hugginggpt · 2026-10-04 · 7624cf3

Defense-in-depth score

2.9 / 10

Minimal

HuggingGPT as shipped runs every request on the operator's OpenAI and Hugging Face credentials, and its default network exposure, credential handling, file handling and web client are not locked down. The model's plan runs without approval and fetches any URL it chooses. Treat it as research demo code: run it only on localhost, with throwaway keys.

Key gaps (3)

  1. The server forwards client-supplied API keys to the downstream LLM API (token passthrough). C1 · Identity & least privilege
  2. A hijacked run can exfiltrate unattended via server-side URL fetches and has further unattended egress paths. C5 · Untrusted input blast radius
  3. Unpinned third-party model files are unpickled with torch.load inside the model server process that holds the config with API keys. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.05 / 1.00

Every request to the HuggingGPT server runs on the operator's own OpenAI key and Hugging Face token, read from the config file or environment at startup. The HTTP API has no authentication, and its default network exposure is not locked down. A request may also supply its own API key, which the server forwards to the model call. There is no per-request authorization anywhere.

C2 Approval gates

Minimal 0.05 / 1.00

HuggingGPT has no approval step at all. The model's task plan is executed immediately: it fetches whatever image or audio URLs the plan names, calls remote or local models, writes generated files into the publicly served folder, and spends the operator's API credit, with no human seeing the plan first. The consequential actions are mostly low-impact (new media files, inference calls), but outbound requests and spend cannot be undone.

C3 Tool & action scoping

Minimal 0.25 / 1.00

The tool set is a fixed list of AI task types, and the code rejects task names outside its model catalogue, which is the main real restriction. Arguments are not validated: image and audio inputs are any URL the model writes, fetched server-side with no host allowlist (internal addresses and cloud metadata included), and local-path handling is not a strict boundary. The model's chosen model id is also used without checking it against the candidate list.

C4 Code-execution isolation

N/A · full credit 1.00 / 1.00

HuggingGPT never interprets model-generated text as code. The model's output is parsed as JSON and only selects task types, model ids and media URLs; the single shell call in the model server runs ffmpeg on a temporary file path the code itself generates. This is a design absence, not a sandbox. Note that the separate EasyTool research scripts in the same repository do pass model output to Python eval(), which is outside the scored component.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Nothing limits what a manipulated model can do. Model results from remote Hugging Face endpoints and content derived from user-supplied media are pasted into the final prompt with the same standing as the user's request. The bundled web client's handling of model output is not locked down, which gives a hijacked response further unattended egress paths. Because the server is shared, others can drive it with the operator's credentials.

C6 Memory, context & configuration integrity

N/A · full credit 1.00 / 1.00

The HuggingGPT server keeps no memory between requests: the conversation history is sent by the client with each call, prompts and demos come only from the operator's --config file, and the logs it writes are never read back into the model. There is no .env loading or workspace instruction file. One caveat outside the scored server mode: the Gradio demo keeps a single global conversation list and API key shared by every visitor.

C7 Third-party extensions

Minimal 0.05 / 1.00

In the default hybrid mode, the local model server loads about thirty third-party model repositories that the setup script clones from Hugging Face at whatever their latest revision is, with no pinning or hash check. At least one checkpoint is loaded with torch.load, which unpickles and can run arbitrary code, and other weights load in the same process. That process runs as the operator and reads the same config file that holds the API keys, so a tampered model file gets everything.

C8 Secrets & sensitive-data protection

Minimal 0.00 / 1.00

The README tells users to paste their OpenAI key and Hugging Face token into the tracked YAML config, with environment variables as an alternative; nothing is masked anywhere. A debug log file at full verbosity is on by default and records every prompt and model response. Outbound credential handling is not locked down. The web client stores the user's key in browser localStorage.

C9 Audit & traceability

Minimal 0.30 / 1.00

The server appends a JSON line per request to log files with the input, task plan, model choices, results and response, plus a full-verbosity debug log. The JSON record has no timestamp, no caller identity, and is only written at the end of a successful run, and the /tasks and /results endpoints, which also execute work, return before any success record is written. Logs live in the server's working directory where the agent has no tool to edit them.

C10 Limits & kill switch

Minimal 0.20 / 1.00

The pipeline has a fixed shape (plan, choose, execute, respond), and a dependency-wait loop gives up after about 80 seconds, but nothing bounds how many tasks the model plans, each of which gets its own thread and paid model calls. Most outbound requests have no timeout, there is no cost cap, no rate limit, and no way to cancel a running request. Nothing limits how many runs callers start on the operator's key.