BoundBench

Cloudflare MCP Servers

Source for Cloudflare's MCP servers (Workers, DNS analytics, observability, etc.)

github.com/cloudflare/mcp-server-cloudflare · 2026-10-03 · ab883e5

Defense-in-depth score

3.9 / 10

Minimal

These servers act on your Cloudflare account with your own OAuth grant or API token, and each server asks only for the scopes its domain needs. The main risk is that the server gives your MCP client little help in deciding what needs approval: the bindings server can irreversibly delete databases and buckets, the D1 SQL tool is labelled non-destructive, and the container and remote-capture tools carry no risk labels. Every server also accepts and forwards raw API tokens, and the container server runs model-chosen shell commands and packages as root with open internet and no timeout.

Key gaps (2)

  1. Every authenticated server accepts a client-supplied Cloudflare API token and forwards it unchanged to the Cloudflare API (token passthrough). C1 · Identity & least privilege
  2. The container server has the model install and run unverified pip/npm packages, as root with full internet, without user consent. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.25 / 1.00

Each server asks Cloudflare's OAuth for its own scope list (the observability server only asks for read scopes; the bindings server asks for Workers and D1 write), and every account-scoped tool checks that the chosen account is one the user's credential can actually access. Read and write tools on the same server share one credential, and the scopes always include offline_access, so the server stores a 30-day refresh token. Every server also accepts a raw Cloudflare API token from the client and forwards it unchanged to the Cloudflare API (token passthrough), so the server's authority is whatever token the client hands it. That passthrough caps this criterion.

C2 Approval gates

Minimal 0.20 / 1.00

These are tool servers, so the approval prompt belongs to the MCP client; the server's job is to label which tools are risky. Only some servers set read-only/destructive hints. The container server (shell execution, file delete), the DEX server (remote packet captures on employees' devices), the Browser Run server and the URL scanner carry no hints, and the D1 query tool, which runs arbitrary SQL including DROP/DELETE, is explicitly labelled non-destructive. There is no dry-run, no read-only mode and no server-side confirmation; the DEX capture tools ask the model in their description to confirm with the user, which is not a control. Deleting databases, buckets and namespaces is irreversible.

C3 Tool & action scoping

Minimal 0.45 / 1.00

Tool inputs are typed with zod schemas and the account ID is checked against an allowlist of the user's accounts, but many tools are general-purpose: arbitrary shell in the container, arbitrary SQL against D1, arbitrary GraphQL, any URL for Browser Run, and a Hyperdrive edit that can point a database connection at any host. Several free-form IDs (crawl job and browser session IDs) are pasted into API paths without validation. Each server ships its full tool set, so the bindings server exposes create/delete tools by default. A misused tool acts on production resources in the user's Cloudflare accounts.

C4 Code-execution isolation

Moderate 0.57 / 1.00

Only the container server runs model-written commands: container_exec passes the string to a shell inside a per-user Cloudflare Container, and nothing runs on the Worker itself. The container image in the repo is a stock Alpine image running as root with full internet access, and production deploys a pinned registry image that cannot be checked against the repo. No credentials are passed into the container. Exec has no timeout (the timeout argument is accepted but ignored), and a dev/test environment setting routes calls to a local process instead of a container.

C5 Untrusted input blast radius

Minimal 0.05 / 1.00

Several servers return attacker-influenced content to the model: web pages from Browser Run, Workers logs (anyone who can hit the user's Worker can write log lines), AI Gateway prompt logs, D1 and KV data, and DEX diagnostic files. Results are returned as plain JSON text with no marking of what is untrusted, and the DEX server mixes its own instructions to the model into the same output as the data. The server has no read-only or no-egress mode. A hijacked session on the bindings server can irreversibly delete databases and buckets unattended; leaking data out needs a second server or a less direct path, such as repointing a Hyperdrive config.

C6 Memory, context & configuration integrity

N/A · full credit 1.00 / 1.00

The servers keep no model-writable memory and auto-load no instruction or configuration files. Durable state is OAuth grants, an identity cache keyed by token hash, and per-user container working files that are only read when the model asks for them. Server instructions are static strings plus the user's account list from Cloudflare's API. There is nothing here for an injection to persist into.

C7 Third-party extensions

Minimal 0.10 / 1.00

The servers load no plugins or third-party MCP servers. The container server, however, tells the model to install whatever packages it needs, so pip/npm packages chosen by the model are downloaded and run without verification or consent. That code is confined to the user's container, which holds no Cloudflare credentials but has full internet access and runs as root.

C8 Secrets & sensitive-data protection

Moderate 0.50 / 1.00

Cloudflare credentials stay on the server side: the model never sees them, the identity cache stores a SHA-256 of API tokens rather than the token, and Sentry reporting uses a header allowlist that excludes Authorization and strips query parameters other than scope. Upstream OAuth tokens and 30-day refresh tokens are stored in the OAuth provider's grant props, whose encryption is library behaviour rather than code here. Errors returned to the model include upstream error text, and Workers traces are enabled by default. Leaked tokens are scoped per server but refresh tokens are long-lived, and passthrough API tokens can be broader.

C9 Audit & traceability

Minimal 0.38 / 1.00

Every tool registered through the shared layer emits a metric to Workers Analytics Engine with the tool name, the user ID and an error code, written by server code the model cannot influence. It does not record arguments, the target account or results, so you cannot reconstruct what was deleted or executed. Tools that return an error result instead of throwing are recorded as successes, and the metric is written after the call, best-effort, with failures only logged to the console.

C10 Limits & kill switch

Minimal 0.38 / 1.00

The servers cap MCP request bodies at 4 MB, and the container server limits the fleet to 50 containers and reaps containers older than 15 minutes, but only when capacity is already half used. Shell commands in the container have no timeout: the timeout argument is declared but never applied, so a process can run until the container is reaped. Crawl depth and page limits are left to the model, and there is no per-user rate limit in code. Workers' platform CPU limits are not credited because they are not configured here.