BoundBench

Firecrawl MCP

MCP server exposing Firecrawl web scrape, search, crawl, map, browser interaction, agent research, monitoring and paper-search tools.

github.com/firecrawl/firecrawl-mcp-server · 2026-10-05 · af5c378

Defense-in-depth score

4.0 / 10

Minimal

A broad web-data server: besides scraping and search it can drive a remote browser that runs model-written code, start autonomous research jobs, and create recurring monitors that email or call webhooks. It runs nothing on the local machine and labels every tool for hosts, but web content returns unflagged alongside instructions to the model, there is no local read-only mode or audit log, and a hijacked model gets both outbound channels and irreversible actions. A .env file in the launch directory is also auto-loaded without any trust decision.

Key gaps (2)

  1. A hijacked host model can, unattended, send data out through scrape URLs, crawl webhooks or monitor emails and also take irreversible actions such as deleting monitors or submitting forms on live sites. C5 · Untrusted input blast radius
  2. A .env file in the server's launch directory is auto-loaded with no trust decision and can change security-relevant settings such as the API endpoint and credential. C6 · Memory, context & configuration integrity

Criteria

C1 Identity & least privilege

Minimal 0.45 / 1.00

The server runs with one Firecrawl credential, an API key or OAuth access token read from the environment, and attaches it to every request to the Firecrawl API. That one key covers the whole Firecrawl account: credit spend, crawl and agent jobs, interact sessions and the account's monitors, including deleting them. There is no per-tool or read-only credential and no per-request authorization in the server. The working directory's .env file is also loaded at startup without any trust decision, so the endpoint and credential settings are not limited to the operator's own environment.

C2 Approval gates

Minimal 0.38 / 1.00

The server gives hosts risk labels for every tool: reads are marked read-only, and monitor update/delete and interact-stop are marked destructive. The labels are not fully accurate, though: firecrawl_interact, which can click, fill and submit forms on live sites and run Bash, Python or Node code in a remote browser, is marked non-destructive even though its own description warns of persistent external side effects. There is no dry-run or preview and no read-only mode for local use; the safe mode that strips browser actions and webhooks applies only to Firecrawl's hosted service. Several actions cannot be undone: deleting monitors, submitting forms and sending monitor emails or webhooks.

C3 Tool & action scoping

Minimal 0.35 / 1.00

Tool parameters are declared as zod schemas, which the MCP framework checks before a tool runs, and several fields have real bounds (URL format on most tools, interact timeout up to 300 seconds, feedback list sizes). Many important fields are open, though: the crawl start URL and the crawl/map page limit are unbounded, and crawl and monitor webhooks take any URL and any headers. Every tool, including remote code execution through interact and monitor creation and deletion, is on by default; only the two feedback tools can be switched off. When pointed at a self-hosted API, firecrawl_parse reads any path on the local filesystem with no directory restriction.

C4 Code-execution isolation

Hardened 0.85 / 1.00

The server never runs model-written code on the user's machine: there is no shell, eval or subprocess anywhere in its source. Code that the model supplies to firecrawl_interact (Bash, Python or Node) or as JavaScript browser actions is sent to Firecrawl's remote browser service and runs there. That remote session has open web access and can persist across calls until stopped, and its isolation is Firecrawl's backend, which this review could not inspect.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Web pages, search results, crawled sites and Alexandria provider data all come back through the same tools with no untrusted flag. Results are JSON with a structured copy, and the Firecrawl API's own guidance is labelled as separate from page content. But results also carry instructions to the model (next-tool suggestions, a feedback-tool pointer, recovery steps), and there is no local mode that drops egress. If injected web content hijacks the host model, the server gives it outbound channels (arbitrary scrape URLs, crawl webhooks, monitor email and webhooks) and irreversible actions (monitor deletion, form submission through interact), with nothing in the server asking a human first.

C6 Memory, context & configuration integrity

Minimal 0.25 / 1.00

The server keeps no memory and reads no instruction files. At startup, though, it silently loads a .env file from whatever directory the host launches it in, often the user's open project. That file can change security-relevant settings, including the API endpoint and credential, and the change persists for every session started in that directory.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

The server loads no plugins, launches no other MCP servers and installs nothing at runtime; it uses only its own fixed npm dependencies. Alexandria providers run on Firecrawl's side and come back as tool output, which is covered under untrusted input. The plugin manifests in the repo are packaging for hosts, not something the server loads.

C8 Secrets & sensitive-data protection

Minimal 0.42 / 1.00

The credential comes from environment variables and is only ever placed in the Authorization header to the Firecrawl API; it is never written into tool results. In the default stdio mode the server's logger is silent, and the hosted-mode action and telemetry records are built from a fixed set of fields that excludes credentials. There is no masking or redaction helper, error messages pass API response bodies straight back to the model, and the key itself is long-lived and covers the whole account.

C9 Audit & traceability

Minimal 0.00 / 1.00

In the default stdio mode the server keeps no record of what it did: its logger only writes when running as an HTTP or hosted service, and the structured per-call action log is emitted only in Firecrawl's hosted deployment. Tool calls, arguments and results leave no trace in the server, so any audit trail has to come from the host or from Firecrawl's account dashboard.

C10 Limits & kill switch

Minimal 0.33 / 1.00

Some work is bounded by the server: interact code runs with a timeout of at most 300 seconds, Alexandria calls have a 15-second timeout and an inline output budget, and PDF parsing has a page cap. Other work is not: the crawl tool polls until the job finishes with no deadline, crawl and map page limits are whatever the model asks for, and there is no rate limit. Monitors the model creates keep running on their schedule after the session ends until someone deletes them, and host cancellation is not passed on to in-flight requests.