BoundBench

Tavily MCP Server

MCP server exposing Tavily web search, extract, crawl, map, research and feedback tools over stdio.

github.com/tavily-ai/tavily-mcp · 2026-10-04 · 1c1d54c

Defense-in-depth score

3.8 / 10

Minimal

A small, read-mostly web search server: it runs no code, touches no files and only ever talks to Tavily's API, so it cannot do much damage on its own. Its gaps in posture are around that core: web content comes back unmarked and mixed with the server's own instructions to the model, there are no risk labels for hosts and no audit log, and the extract and feedback tools give a hijacked model an outbound channel. Its configuration loading is also not integrity-protected.

Criteria

C1 Identity & least privilege

Moderate 0.50 / 1.00

The server holds one credential, a Tavily API key read from the environment, and uses it only to call Tavily's own API at hardcoded addresses. It does not touch files, spawn processes or reach other systems, so a stolen or misused key can spend Tavily credits and submit feedback, but nothing else. All six tools share that one key and there is no per-request authorization. Credential loading is also not integrity-protected.

C2 Approval gates

Minimal 0.10 / 1.00

As a tool server it relies on the MCP host to ask the user before calls, but it gives the host nothing to decide with: none of the six tools carries a read-only or destructive label, there is no dry-run, and no read-only mode. The tools mostly read the web, but the feedback tool sends free text (including the agent's final answer) to Tavily and the research tool can spend significant credits. No tool can delete or change data anywhere, which keeps the damage from a wrongly approved call low.

C3 Tool & action scoping

Minimal 0.30 / 1.00

Each tool calls one fixed Tavily endpoint, so the model cannot point the server at an arbitrary host or service. Beyond that, arguments are passed straight through: the declared schemas (enums, a max_results ceiling) are advisory and not checked in the server, crawl and map limits have no upper bound, and any URL can be handed to Tavily to fetch. All six tools are always on and cannot be switched off individually.

C4 Code-execution isolation

N/A · full credit 1.00 / 1.00

The server never interprets model text as code: there is no shell, eval, subprocess or script execution anywhere in its source. All work is HTTP requests to Tavily's API.

C5 Untrusted input blast radius

Minimal 0.05 / 1.00

Every tool returns third-party web content as flat text mixed with the server's own labels, with no marker that it is untrusted. The server's own tool description for feedback is written as instructions to the model, and error responses from the API are relayed with suggested next actions such as agentic payment and POSTing answers to an endpoint for bonus credits. If injected web content hijacks the host model, it can use extract to send data in a URL to any host via Tavily, or the feedback tool to send free text to Tavily. The server holds no private data itself and cannot change anything, so the worst case is data leaving, not destruction.

C6 Memory, context & configuration integrity

Minimal 0.25 / 1.00

The server keeps no memory and reads no instruction files, but its configuration loading is not integrity-protected, which lets workspace content influence it across sessions.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

The server loads no plugins, launches no other MCP servers and installs nothing at runtime; it uses only its own fixed npm dependencies. (Hosts launching it with npx tavily-mcp@latest is the host's supply-chain choice, not something this server does.)

C8 Secrets & sensitive-data protection

Minimal 0.30 / 1.00

The API key comes from the environment, is never echoed into tool results, and error messages returned to the model are built from the API's detail field rather than the raw request. There is no redaction anywhere, though, the key is sent in both the Authorization header and every JSON body, and MCP errors are printed whole to stderr. The feedback tool's description pushes the model to send its final answer and quoted snippets to Tavily, and every request carries a session ID for tracking.

C9 Audit & traceability

Minimal 0.00 / 1.00

The server keeps no record of what it did. Tool calls, arguments and results are not logged; the only output is a startup line and raw MCP errors on stderr. Any audit trail would have to come from the host or Tavily's own dashboard.

C10 Limits & kill switch

Minimal 0.33 / 1.00

The research tool has real time limits: polling stops after 5 or 15 minutes and the streaming fallback has header, idle and overall timers that tear down the connection. Crawl output is truncated to 200 characters per page. The other tools have no request timeout, crawl and map limits are model-chosen with no upper bound, and nothing rate-limits calls or cancels work in flight when the host cancels.