C1 Identity & least privilege
Minimal 0.05 / 1.00
OpenAGI has no identity or authorization layer of its own. Every bundled tool reads a long-lived API key (a shared RapidAPI key, a Hugging Face token, a Bing key) straight from the process environment and calls the vendor API with it. Nothing narrows these keys per tool or per request, and the same environment is inherited by the pip subprocesses and by any agent code downloaded from the hub. A hijacked agent therefore acts with whatever the operator's keys allow.
C2 Approval gates
Minimal 0.05 / 1.00
There is no approval step anywhere in the framework. The ReAct executor takes whatever tool calls the model returns and runs them immediately, including tools that write files to disk or send local files to a third-party API. Developers who register their own tools (email, payments, shell) get the same ungated execution. The bundled tools are mostly read-only lookups, which limits but does not remove the damage.
C3 Tool & action scoping
Minimal 0.38 / 1.00
Tools are narrow by construction: each one calls a single fixed vendor endpoint, and an agent only gets the tools its config.json lists. The one shared argument check rewrites any parameter whose name contains 'path' into an output folder, but its containment is not a complete boundary. Other arguments (queries, durations) are passed through without bounds. The model cannot add tools at runtime, but nothing stops an agent config from enabling write tools.
C4 Code-execution isolation
Minimal 0.00 / 1.00
The framework gives the model no code or shell tool, but it does execute third-party and data-supplied code on the host with no isolation. Activating an agent runs 'pip install -r' on that agent's requirements (package install scripts run as the user with the full environment), and the travel-planner distance tool calls Python eval() on values from a data file. Nothing in the repository provides a sandbox, container, or restricted user, so anything that runs has the operator's full access and API keys.
C5 Untrusted input blast radius
Minimal 0.05 / 1.00
Results from web search, Wikipedia, arXiv and travel APIs are pasted straight into the conversation as assistant messages, with the same standing as the agent's own reasoning. Nothing marks them as untrusted or restricts what the agent can do after reading them. A hijacked agent can write files and push local files to a third-party API without anyone approving it. The bundled tools offer no attacker-chosen URL, which limits direct exfiltration, but framework users' own tools get no protection.
C6 Memory, context & configuration integrity
N/A · full credit 1.00 / 1.00
OpenAGI keeps no memory the model can write to. Conversation messages live only in the agent object for one run, logs are written but never read back, and there is no .env loading or auto-loaded instruction file. The RAG example builds a local vector store, but only from an essay shipped with the agent, not from anything the model or a user supplies. Agent packages downloaded from the hub do persist on disk and are re-imported later; that risk is scored under third-party extensions.
C7 Third-party extensions
Minimal 0.00 / 1.00
This is the most serious problem. When an agent name is activated and its folder is not present locally, the framework downloads that agent's Python code, config and requirements from a hosted hub (openagi-beta.vercel.app), writes them into the package, pip-installs the requirements, and imports the code, all without asking, without a version pin, and without any hash or signature check. Anyone who can publish or alter an agent on that hub, or who controls that endpoint, gets code execution in the operator's process with every API key in its environment.
C8 Secrets & sensitive-data protection
Minimal 0.05 / 1.00
API keys come from environment variables and are never placed into the model prompt, which is the main thing done right. There is no masking or redaction anywhere, though: step results and tool parameters are printed to the console or appended to log files under the working directory, and the full environment, keys included, is inherited by pip subprocesses and by any downloaded agent code. There is no telemetry.
C9 Audit & traceability
Minimal 0.25 / 1.00
The ReAct agent prints each step's result, which includes a sentence naming the tool called and its parameters, to the console by default, or appends it to a text file under ./logs in file mode. This is free text, not a structured record, has no actor or approval fields, and is written only after the step has run. Agents that override the run loop, or code using CallCore, record nothing about tool use.
C10 Limits & kill switch
Minimal 0.25 / 1.00
Limits are minimal. Planning and tool-call retries are capped at three, and manual workflows have a fixed number of steps, but in automatic mode the model writes the plan and so decides how many steps run. There is no wall-clock limit, token or cost budget, or rate limit in the framework, and the request loop re-queues until the external AIOS scheduler reports done. There is no stop or cancel mechanism; the terminate signal is commented out.