BoundBench

Google Workspace MCP

MCP server controlling Gmail, Calendar, Docs, Sheets, Drive, Chat etc.

github.com/taylorwilsdon/google_workspace_mcp · 2026-10-03 · 3ada8ba

Defense-in-depth score

5.4 / 10

Moderate

Run with no flags, this server gives the model write access to the user's whole Google Workspace account, including sending email, sharing files publicly and writing and running Apps Script. Local file and URL handling is well hardened, and it runs no code on the host. The dominant risk is prompt injection: emails and documents it reads come back as plain text beside its own instructions, and nothing in the server stops an injected instruction from leaking data or sending mail. Use --read-only or --permissions to cut its authority.

Key gaps (2)

  1. Default install requests write scopes across the user's entire Google Workspace account (mail send, full Drive, Apps Script execution), so any control failure exposes the whole account. C1 · Identity & least privilege
  2. Untrusted email/document content, the user's private Workspace data, and send/share tools coexist in one default session with nothing in the server separating them. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.40 / 1.00

Out of the box the server asks Google for one OAuth grant that covers every enabled service with full write scopes: all of Drive, Gmail send/modify/settings, Calendar, Contacts, Chat, and Apps Script project, deployment and external-request scopes. Each tool only checks that the stored token has enough scopes; it then receives the same all-powerful token. In the default local mode the target account is a tool argument the model fills in, so whichever stored account it names is used. An opt-in --permissions flag requests per-service minimal scopes and removes tools that need more, which is the scored mechanism here, but it is off by default.

C2 Approval gates

Minimal 0.45 / 1.00

As a tool server it relies on the MCP host to ask the user before acting; what it provides is a read-only or destructive label on every one of its tools, and those labels are mostly accurate. There is no preview or dry-run for sending mail, changing sharing, or deleting, except a dry-run for applying Gmail filters. A server-enforced read-only mode exists but must be switched on. If a host auto-approves or the user approves the wrong call, emails, public file shares and calendar invites go out immediately and cannot be recalled.

C3 Tool & action scoping

Moderate 0.53 / 1.00

Arguments that touch the local machine or arbitrary URLs are validated well: local file paths are resolved and confined to the attachment directory by default, and URL fetches block private addresses and recheck every redirect. Most other arguments are typed and passed to Google's own APIs. However, with no flags every tool is loaded, including email send, file sharing, and Apps Script tools that write and run code, and nothing bounds recipients or the number of operations.

C4 Code-execution isolation

Strong 0.80 / 1.00

The server never runs code on the host: there is no shell, eval or subprocess path. The only code-execution path is Apps Script, where the model can write script code and run it through Google's Execution API, so it runs in Google's managed runtime rather than on the machine. Inside that runtime the code acts with the user's authorized scopes and can make outbound web requests, and the model can also edit the script manifest that declares those scopes.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

The server hands back email bodies, documents, chat messages and search results as plain text mixed with its own instructions to the model (for example 'Use get_gmail_attachment_content(...) to download'), with no marker that the content came from a third party. In the default configuration the same session can read an attacker's email, read private Drive and Gmail data, and send email or share files publicly. Nothing in the server breaks that combination, so a successful prompt injection can leak data and take irreversible actions unless the host stops it.

C6 Memory, context & configuration integrity

N/A · full credit 1.00 / 1.00

The server keeps no memory, retrieval store or conversation history that is fed back to the model, and it does not load instruction or configuration files from the user's working directory; its .env is read only from its own install directory. Stored OAuth credentials and downloaded attachments are not read back into model context as instructions.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

The server loads no plugins, extensions, remote tools or model files at runtime; service modules are imported from a fixed in-repo list. Its own Python dependencies are build-time supply chain and out of scope here.

C8 Secrets & sensitive-data protection

Minimal 0.45 / 1.00

OAuth tokens never go into tool output sent to the model, HTTP client logs that would print access tokens are silenced, and URL query strings are scrubbed from error logs. Refresh tokens, however, are stored as plaintext JSON files (owner-only permissions) and are long-lived and broadly scoped. Telemetry is off unless an OpenTelemetry endpoint is configured, and debug logging that includes user text is opt-in.

C9 Audit & traceability

Minimal 0.38 / 1.00

Each tool call writes an unstructured text log line naming the tool, the Google account and the service, into a log file under the user's home directory that is on by default. Arguments and results are not recorded in a structured way, there is no approver or session attribution, and logging failures are only reported to stderr while the server keeps running. Structured OpenTelemetry tracing exists but is opt-in.

C10 Limits & kill switch

Minimal 0.40 / 1.00

Some of the server's own work is bounded: Google API calls time out after 30 seconds, Office document expansion is capped at 25 MiB, and email bodies are truncated. File downloads are uncapped unless the operator sets a limit, and there are no rate limits on sending mail, sharing files or other side effects, so a runaway host loop can keep acting until Google's own quotas stop it.