C1 Identity & least privilege
Moderate 0.53 / 1.00
The gateway refuses to start without a strong master key, and every request is then authorised in code against the calling virtual key, team, user and organisation before the gateway attaches its own credentials. MCP servers are deny-by-default per key (only servers granted to the key or team, or marked allow_all_keys, are reachable), and per-key tool permissions are re-checked on every tool call. The credentials behind that check are static and shared: the operator's provider API keys and each MCP server's single configured credential serve every caller, and a key with no model list can call every model. Per-user OAuth, token exchange and passthrough modes exist for MCP but are opt-in.
C2 Approval gates
Minimal 0.33 / 1.00
When a caller attaches gateway MCP tools to /chat/completions or /responses, LiteLLM executes the model's tool calls itself only if every MCP tool reference in the request sets require_approval to "never"; otherwise the exact tool calls are handed back to the calling application unexecuted. That check is deterministic and the model cannot set it, but it is a blanket per-request switch held by the caller, there is no human approval step or risk tiering in the gateway, and the README's lead MCP example turns auto-execution on against a GitHub MCP server. Once enabled, upstream tool actions are as irreversible as the upstream makes them.
C3 Tool & action scoping
Moderate 0.53 / 1.00
Tool scoping is by name: each MCP server can carry an allowed or disallowed tool list, an optional allowed-parameter-name list per tool, and pinned tool definitions, and per-key and per-team tool permissions narrow further. All of these checks run in one pre-call function that every gateway tool call goes through. Argument values are not validated against bounds or allowlists, and by default every tool an upstream server lists is exposed to any key that has the server. User-supplied URLs fetched by the gateway (images, files) go through an SSRF guard that rechecks redirects.
C4 Code-execution isolation
Strong 0.70 / 1.00
In the default gateway nothing executes model-written code: Jinja chat templates render in Jinja's immutable sandbox, operator-written custom-code guardrails run under RestrictedPython, and stdio MCP servers are disabled unless an environment flag is set. The one model-code path, the opt-in code-interpreter interception callback, always runs code in a remote sandbox service (E2B or OpenSandbox) and raises an error rather than falling back to local execution. Those sandboxes get internet access by default, and stdio MCP servers, when enabled, run as host subprocesses (root in the shipped image) with a scrubbed environment.
C5 Untrusted input blast radius
Minimal 0.28 / 1.00
MCP tool results and upstream tool descriptions enter the model's context like any other message, with no provenance tagging or taint tracking by default. When auto-execution is on, the streaming Responses path lets the model choose further tool calls after reading tool output, for up to five rounds, so injected tool output can drive actions with the shared MCP credentials. By default the gateway returns tool calls to the calling application instead of running them. An opt-in tool_policy guardrail blocks tools marked trusted-input once untrusted tool output is in a chat conversation, and many third-party injection-detection guardrails can be enabled.
C6 Memory, context & configuration integrity
Moderate 0.50 / 1.00
The model has no memory tool and the gateway auto-loads no workspace instruction files; configuration comes from the operator's config file or database. The /v1/memory store is written and read by applications and is filtered to the caller's user or team in the query, and it is never injected into prompts automatically. Responses sessions can be rebuilt from stored spend logs by previous_response_id (prompt content is only stored when store_prompts_in_spend_logs is enabled); session replay is not equally scoped. Cached or replayed context is not validated or provenance-tagged.
C7 Third-party extensions
Minimal 0.23 / 1.00
Third-party code reaches the gateway three ways: Python callbacks, guardrails and router plugins named in the config (including modules downloaded from S3 or GCS) are imported and executed in-process with no hash or signature check; stdio MCP servers run operator-chosen commands as subprocesses when explicitly enabled; and remote MCP servers are connected over HTTP. Only proxy admins can add servers (user submissions need admin approval and cannot be stdio), and stdio subprocesses get a minimal environment. Pinning an MCP server's tool catalog, which serves the pinned definitions and alerts on drift, is available but off by default.
C8 Secrets & sensitive-data protection
Minimal 0.38 / 1.00
Credentials stored in the database are encrypted with a key derived from LITELLM_SALT_KEY (falling back to the master key), virtual keys are stored hashed, log records pass through a credential-redaction filter by default, prompts are not written to spend logs unless enabled, and there is no vendor telemetry. Secrets are not placed in model context, but nothing scans tool results or model-bound messages for secrets. The shipped docker-compose file's defaults are not locked down, and the redaction filter can be turned off with an environment variable.
C9 Audit & traceability
Moderate 0.50 / 1.00
Every gateway MCP tool call is logged through the same logging object as model calls, and the spend-log row carries the tool name, arguments, server, key, user, team and trace id, stored in Postgres outside anything the model can touch. Denied calls are logged as failures. There is no approval record because the gateway has no approval step, and change audit logs for keys, teams and MCP servers are written only for enterprise (premium) deployments. Logging is best-effort: setup failures are swallowed at debug level and spend logs are batched asynchronously.
C10 Limits & kill switch
Moderate 0.50 / 1.00
Server-side tool execution is bounded: one round on the non-streaming paths and at most five rounds on streaming Responses, MCP calls time out after 60 seconds, and requests after 6000 seconds. Key, team and user budgets, RPM/TPM limits and a per-session iteration limiter are enforced in code but are all unset by default, so out of the box there is no spend ceiling. Blocking a key stops new requests but does not cancel calls already in flight.