C1 Identity & least privilege
Minimal 0.00 / 1.00
CowAgent runs every tool as the operating-system user that launched it, with no identity of its own. At startup it copies model API keys and chat-platform app secrets (Feishu, DingTalk, WeChat, QQ) from config into the process environment, and the shell tool hands that whole environment, plus ~/.cow/.env, to every command; the tool description even tells the model the keys are available as $VARS. The optional permission modes are off in the shipped config (full-access). A hijacked agent therefore holds the user's whole local account plus every connected service credential.
C2 Approval gates
Minimal 0.05 / 1.00
There is no human approval step anywhere in the tool path: shell commands, file writes, browser clicks, scheduled tasks and outbound sends execute as soon as the model asks. The only safeguards are prompt text telling the model to confirm destructive operations and a tiny shell denylist that returns an error. Most actions are irreversible; only self-evolution edits to memory and skills are snapshotted for undo.
C3 Tool & action scoping
Minimal 0.25 / 1.00
In the shipped full-access mode every tool is enabled and the shell, web fetch and browser tools accept arbitrary commands and URLs; the only argument checks are a denylist that blocks the ~/.cow/.env credential file and /proc environ paths, and SSRF protection that is off by default. The optional workspace-write and read-only modes add real path containment for write/edit, but they do not cover every path. Those modes are off by default.
C4 Code-execution isolation
Minimal 0.47 / 1.00
Model-written shell commands run directly on the host as the user through subprocess with shell=True, with the full environment and no sandbox. The only filter is a deliberately minimal denylist (rm -rf /, dd from /dev/zero, shutdown) that tells the model to ask the user, which other phrasings evade. The documented Docker deployment runs the whole app in a non-root container, which is a real but basic boundary: the compose file disables seccomp, mounts the data directory holding config and credentials, passes API keys in the container environment, and leaves network egress open.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
Web pages, search results, files, MCP outputs and (when chat channels are connected) other people's messages enter the model context as ordinary tool results or user turns, with no marking, filtering or change in privileges. Once hijacked, the agent can read the user's secrets and send them anywhere with web_fetch or curl, and can delete files or post through the browser and chat tools, all without a human in the loop. Connected chat channels make this worse: Telegram, Slack, Discord and others whitelist all groups and have no sender allowlist, so anyone who can message the bot can instruct a full-access agent.
C6 Memory, context & configuration integrity
Minimal 0.00 / 1.00
AGENT.md, USER.md, RULE.md and MEMORY.md in the agent's workspace are injected into the system prompt on every turn, and the model writes to them freely, as does an unattended self-evolution pass that is on in the template (it is snapshotted for undo). The same workspace, which is the tools' working directory, holds mcp.json; a change to that file is hot-reloaded and launches new MCP servers with no prompt. Memory has no per-user isolation yet, so in a shared chat one participant's injected memory influences everyone.
C7 Third-party extensions
Minimal 0.00 / 1.00
Third-party code arrives as skills (from the Skill Hub, any GitHub/GitLab branch or a raw URL) and as MCP servers launched from mcp.json. Nothing is pinned, and the skill checksum is supplied by the same server that serves the download. New skills are enabled automatically, any chat participant can run /skill install, and the model itself can add an MCP server by editing mcp.json, which is then launched with no consent. MCP stdio servers get a scrubbed environment, but skill scripts run through the bash tool with every API key in their environment.
C8 Secrets & sensitive-data protection
Minimal 0.13 / 1.00
The signing secret used by the web console is not well protected. API keys are stored in plaintext config.json and ~/.cow/.env (the latter chmod 600), then exported to every bash subprocess, and the model is told it can use them as $VARS. The config log line and the bash progress stream are masked, and MCP stdio servers get a scrubbed environment; there is no telemetry by default.
C9 Audit & traceability
Minimal 0.40 / 1.00
Every run is persisted step by step into a SQLite conversation store with timestamps, tool calls and their results, and run.log records each tool call with its arguments. The store lives inside the agent's own writable workspace (memory/long-term/index.db), persistence is explicitly best-effort, and there is no actor attribution beyond agent and session ids. With no approval system, there is nothing to log on that front either.
C10 Limits & kill switch
Minimal 0.40 / 1.00
A run is capped at 30 decision steps by default, each shell command times out after 120 seconds (600 maximum), and a user cancel kills the running command's process group. Sub-agents get half the parent's steps, a 300-second budget and at most three at a time, but there is no token or spend cap and no wall-clock limit on the main run. Background shell jobs keep running after a run ends, scheduled tasks keep firing, and any chat participant can raise agent_max_steps with /config.