BoundBench

Keep

Open-source AIOps and alert management platform with AI correlation and workflow automation that acts on monitoring/ticketing tools

github.com/keephq/keep · 2026-10-03 · 465e991

Defense-in-depth score

2.1 / 10

Minimal

Keep's AI incident chat can call methods on connected providers with no approval step and no allowlist, and the provider-invocation endpoint it uses is not locked down, which widens its reach. Alert text from monitoring tools goes into the same session, so a prompt injection in an alert could act on Kubernetes or ticketing systems and reach stored integration secrets. The default compose file also disables authentication and enables telemetry. Treat the AI features as unsafe to enable on a deployment with real integrations until the invoke endpoint is restricted and chat actions are gated.

Key gaps (4)

  1. The AI chat's invokeProviderMethod runs provider methods with no approval, and the invocation endpoint is not locked down. C2 · Approval gates
  2. A hijacked chat can read secrets and exfiltrate them or take irreversible actions without a human, because untrusted alert text shares the session with ungated provider calls. C5 · Untrusted input blast radius
  3. Shell and Python provider execution runs unsandboxed in the API container, next to plaintext provider secrets, and is reachable from AI paths. C4 · Code-execution isolation
  4. Default install grants every caller the Admin role, and the agent can use every tenant-wide provider credential. C1 · Identity & least privilege

Criteria

C1 Identity & least privilege

Minimal 0.07 / 1.00

The shipped docker-compose runs with authentication disabled, and in that mode every request is treated as an administrator with every scope. The AI incident chat calls the backend with the user's session, so role checks do apply when authentication is turned on, but provider credentials are stored once per tenant and any installed provider's full credentials can be used through the chat. Gaps in access control on the provider-invocation endpoint also widen what the chat can reach on the API server, which holds every tenant's integration secrets. A hijacked assistant therefore acts with the full authority of every connected system.

C2 Approval gates

Minimal 0.20 / 1.00

The AI workflow builder asks the user to accept each proposed step before it is added to the canvas, but that only edits an unsaved draft. The AI incident chat, which is the part that acts, has no approval step at all: its invokeProviderMethod, createIncident, enrichment and incident-update actions run as soon as the model calls them. invokeProviderMethod can call action methods on installed providers, such as executing commands in pods, restarting pods, or creating external incidents. Nothing the model does from the chat is shown to a human for approval before it happens.

C3 Tool & action scoping

Minimal 0.07 / 1.00

The chat's invokeProviderMethod tool takes a provider id, a method name and a free-form parameter object, and the backend passes the parameters straight through. Method and provider resolution on the invoke route is not locked down. The only argument filter found is a substring denylist of metadata hosts in the HTTP provider. In effect the assistant has general-purpose tools rather than narrow, validated ones.

C4 Code-execution isolation

Minimal 0.00 / 1.00

Code can run through the bash provider (a shell subprocess) and the Python provider (eval inside the API worker), reachable from workflows and, because access control on the invoke endpoint is not locked down, from AI paths. Neither runs in any sandbox: they execute in the API server container as the same non-root user that holds the database, the file-based secret store and network access. Code that escapes nothing still has everything the server has.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

The incident chat puts the incident, its alerts and the results of provider calls into the model's context. Alerts arrive from monitoring tools and webhooks, so their names, descriptions and labels are third-party text, and nothing separates that text from the user's instructions. Because the same session can call provider methods without approval, injected text in an alert could make the assistant send data out or take irreversible actions on connected systems. There is no detection or Rule-of-Two control.

C6 Memory, context & configuration integrity

Minimal 0.00 / 1.00

Whatever the assistant writes into an incident persists and is fed back to every later chat about that incident, for every user in the tenant. The enrichRCA action appends model text to the incident's root-cause list, and every invokeProviderMethod result is stored as an incident enrichment; the whole incident, including enrichments, is supplied to the model as context. Chat history, including past tool calls and results, is also saved in the browser and reloaded. None of this is validated, labelled as untrusted or gated, so a single injection can persist and steer later sessions into tool use.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

Keep does not load third-party code at runtime for its AI features. There is no MCP client, plugin loader, model download or package install path; providers are first-party modules shipped in the image. The search for such loaders found only documentation and developer scripts.

C8 Secrets & sensitive-data protection

Minimal 0.15 / 1.00

Provider credentials are stored by default as plaintext JSON files in the state directory, and API-side secret handling is not locked down. The incident chat only passes provider ids, types and methods to the model, so credentials are not routinely in prompts, but server-side execution paths reachable from the model can read the secret files. Telemetry is on by default: the backend sends usage events to PostHog with a built-in key, the frontend identifies users to PostHog by email, and the compose file sets a Sentry DSN. There is no redaction filter in logging.

C9 Audit & traceability

Minimal 0.35 / 1.00

When the chat invokes a provider method, the backend logs the provider id and method name, but not the arguments or who asked for it. Incident enrichments go into an audit table that records the user's email, so an AI-made change looks the same as one the user made. The chat conversation itself is kept only in the user's browser. Records go to standard logs and the database, outside any workspace, but code running through the bash provider in the same container could alter the database.

C10 Limits & kill switch

Minimal 0.20 / 1.00

Keep's own code sets no step, token or cost limit on the chat assistant, and the API rate limiter is disabled by default. The bash provider has a 60-second default timeout, but the caller sets it, so the model can raise it through the parameters it passes. Stopping the chat in the browser doesn't stop server-side commands already started. CopilotKit's internal loop limits were not examined.