C1 Identity & least privilege
Minimal 0.05 / 1.00
LlamaIndex has no identity or authorization layer for agents: every registered tool runs with whatever credentials the developer's process and the tool's own client hold, and the default LLM reads the operator's OPENAI_API_KEY from the environment. The one binding primitive, FunctionTool partial_params, is not a strict boundary. The framework neither narrows nor widens the developer's authority.
C2 Approval gates
Minimal 0.00 / 1.00
There is no approval gate in the agent loop. The model's tool calls are dispatched straight to the tool function, including CodeActAgent's code-execution tool. Human-in-the-loop is possible only if the developer writes their own wait-for-event logic inside each tool. Nothing in the framework checkpoints or previews actions.
C3 Tool & action scoping
Minimal 0.30 / 1.00
Tool schemas are generated from function signatures and shown to the model, but nothing validates the model's arguments against them before the call: the kwargs go straight into the Python function. The handoff tool is the one built-in that checks its argument against an allowlist of agents. On the positive side, agents start with no tools at all and receive only what the developer lists.
C4 Code-execution isolation
Moderate 0.50 / 1.00
Core ships no isolation for model-generated code. CodeActAgent requires the developer to pass an execute function, and the official example implements it with in-process exec(), labelled not safe for production. The code-interpreter tool package runs model code with subprocess on the host, with no timeout and the full environment. An Azure Dynamic Sessions tool provides a real remote sandbox, but it is a separate opt-in package.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
Tool results, retrieved documents, and reader output enter the model's context as ordinary tool or observation messages with no provenance or taint, and nothing changes what the agent may do after reading them. The project's SECURITY.md explicitly treats prompt injection as the application's problem. Since the framework provides no approval gate either, a hijacked agent can use any registered egress or write tool unattended.
C6 Memory, context & configuration integrity
Minimal 0.28 / 1.00
By default an agent's memory is an in-process chat buffer that disappears with the run, and core loads no workspace files such as .env. When developers enable the richer Memory class with a database and memory blocks, model-extracted facts and retrieved text are written without validation and inserted into the system message by default. Sessions are keyed by an ID in database queries, but nothing prevents poisoned entries from persisting.
C7 Third-party extensions
Minimal 0.05 / 1.00
Extensions (tool packages, MCP servers, rerank and embedding models) are chosen in code by the developer, so nothing is added silently. But nothing is verified: MCP servers launch whatever command is configured, and model loading in one core component is not locked down. Python tool packages also run in-process with all the agent's credentials. One positive: persisted object mappings load through an allowlisted unpickler.
C8 Secrets & sensitive-data protection
Minimal 0.30 / 1.00
Provider keys come from environment variables or plain string fields on the LLM classes. The base LLM class has a to_payload method that keeps API keys out of instrumentation and callback payloads, and core ships no telemetry. There is no redaction of secrets in tool results sent to the model, in logs, or in subprocess environments.
C9 Audit & traceability
Minimal 0.35 / 1.00
Every tool call goes through one workflow step that emits structured ToolCall and ToolCallResult events with the tool name, arguments, call ID, and output, so a developer can capture a full trajectory. But nothing durable records them by default: instrumentation handlers are no-ops, and the run keeps its tool-call list only in memory, returned to the caller. There is no actor or approver attribution.
C10 Limits & kill switch
Minimal 0.38 / 1.00
Agent runs stop after 20 iterations by default, enforced in code, and in multi-agent workflows the counter is shared across handoffs. A wall-clock timeout exists in the underlying workflow engine, but agents pass timeout=None by default, and there is no token or cost cap and no per-tool timeout. Sync tools run in a thread pool, so a cancelled or timed-out run can leave a tool still running.