BoundBench

Composio

SDK monorepo (TypeScript and Python) that gives AI agents 1000+ Composio-hosted toolkits, per-user auth sessions, triggers, and a remote code sandbox.

github.com/composiohq/composio · 2026-10-04 · 98b10d7

Defense-in-depth score

3.4 / 10

Minimal

Composio keeps provider tokens on its servers, never runs model code on your machine, and hardens its own file and URL handling well. But a default session hands the agent every toolkit the user has connected, a generic executor, and a remote sandbox that can call any tool and reach the internet, with no approval step and nothing to separate injected content from instructions. The dominant risk is prompt injection: an email or issue the agent reads can make it leak the user's data and send, post, or delete across their apps unattended. Restrict toolkits and tags, and add approval in your host framework, before production.

Key gaps (2)

  1. Default sessions expose every toolkit with the user's full connected-account authority, so a hijacked agent can act across all of the user's apps. C1 · Identity & least privilege
  2. Default sessions pair untrusted tool output (email, issues, web) with private data and unattended send/delete/egress tools, and search results inject model-directed instructions. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.35 / 1.00

A Composio session is bound to one end user, and every tool call executes server-side with that user's connected accounts, so the agent never holds raw provider tokens. But a default session exposes every toolkit, lets the agent start new connections for more apps, and acts with whatever OAuth scopes the auth config requested, with one credential for reads and writes. The SDK itself authenticates with a long-lived project API key that can act for any user in the project. A hijacked session can therefore act across the user's whole connected footprint.

C2 Approval gates

Minimal 0.15 / 1.00

Composio does not own the agent loop, so the host framework decides what to approve. What Composio gives the host is weak: the default session funnels every app action through one generic executor (COMPOSIO_MULTI_EXECUTE_TOOL, up to 50 tools per call) and a remote code sandbox that can call any tool, so a host can only gate the whole executor, reads included. Risk hints exist on these meta tools, but the default providers don't pass them into the host's approval hooks, and nothing in the SDK pauses for a human. The opt-in TypeSafe provider can require an explicit confirm flag for destructive tools.

C3 Tool & action scoping

Minimal 0.40 / 1.00

The SDK's own local surfaces are carefully validated: automatic file upload/download is off by default, local paths are realpath-checked against an upload directory allowlist and a sensitive-path denylist, and URL fetches go through an SSRF guard that pins DNS, blocks internal and metadata addresses, and rechecks every redirect. But the default tool set is the opposite of narrow: a generic executor for any of 1000+ toolkits, a remote Python/bash sandbox, and an arbitrary-HTTP proxy to connected APIs. Toolkit, tool, and tag filters exist but must be configured.

C4 Code-execution isolation

Moderate 0.57 / 1.00

No model-generated code ever runs on the developer's machine: the SDK has no local exec path, and the default code tools (COMPOSIO_REMOTE_WORKBENCH and COMPOSIO_REMOTE_BASH_TOOL) run in Composio's remote sandbox. The sandbox's internals (container or VM, hardening) are not in this repo, so its strength can only be inferred. By the project's own docs the sandbox is persistent within a session, has full programmatic access to every Composio tool with the user's credentials, can make arbitrary API calls and web searches, and installs packages on demand, so code inside it holds the session's full authority.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Nothing in the SDK limits what injected content can make the agent do. Tool results from emails, issues, chat, web search, and documents go back to the model as plain data with no provenance or untrusted marker, and Composio's own search output adds model-directed instructions (execution guidance, recommended plan steps, server-side tool memory) that the tool description tells the model it MUST follow. A default session combines untrusted input, the user's private data, and unattended irreversible actions and egress, so a successful injection can both leak data and act.

C6 Memory, context & configuration integrity

Minimal 0.35 / 1.00

Neither SDK auto-loads instruction files or a working-directory .env, and the CLI only accepts organization and project IDs from a repo's .composio/.env, explicitly to block credential or endpoint injection from cloned repos. Persistence lives server-side: sessions keep tool memory, sandbox state, and a /mnt/files mount, and the meta-tool search returns stored memory and plans back to the model. How that memory is written, validated, or expired is not in this repo.

C7 Third-party extensions

Minimal 0.25 / 1.00

The agent does not load plugins or third-party code into the developer's process: toolkits are Composio-run integrations executed on its servers, and custom tools are the developer's own code. But tool definitions are fetched at 'latest' by default and the provider execution path skips the version check, so definitions can change under a running agent without re-approval. Per its docs, the remote sandbox installs packages the agent asks for from a supported list, which can't be verified here.

C8 Secrets & sensitive-data protection

Minimal 0.40 / 1.00

Provider OAuth tokens are held server-side and injected at execution time, so they never pass through the SDK or the model. The SDK redacts every log line and the free-form error text it sends in telemetry. Gaps: telemetry is on by default (content-free), the CLI stores the project API key in plaintext unless the keychain is chosen, tool results reach the model unredacted, and the long-lived project API key can act for every user in the project.

C9 Audit & traceability

Moderate 0.50 / 1.00

Every tool execution returns a server-side log ID, and the backend keeps searchable execution logs attributed to user, session, and connected account, outside anything the agent can edit. The logs live in Composio's closed backend, so their completeness and integrity can't be checked here, and there's no record of approvals because nothing is approved. Custom tools that run in the developer's process are only logged if they call back through Composio.

C10 Limits & kill switch

Minimal 0.38 / 1.00

The SDK caps its own work in a few places: URL and file fetches are limited to 100 MiB and five redirects, and the remote bash tool has a stated three-minute limit. There's no limit on how many tool calls, sandbox runs, or bulk executions a session can make, no spend ceiling, and cancelling only stops the SDK waiting, not work already running on the server.