C1 Identity & least privilege
Minimal 0.07 / 1.00
Agents act with the full authority of the user who owns them. Every command call receives all of the agent's stored settings plus every OAuth access token the user has connected, and the user's GitHub token is placed inside the code sandbox. There is no per-tool credential scoping. The built-in 'Create AGiXT Agent' command lets the model create a new agent, enable AI-chosen commands on it from every configured integration, and then enable a delegation command on itself, so the model can widen what it can reach. The default deployment also mounts the Docker socket into the root AGiXT container.
C2 Approval gates
Minimal 0.00 / 1.00
There is no human approval step for any command. Once a command is enabled for the agent (and all Core Abilities are enabled by default), the model's tool calls run immediately: shell commands, Python execution, file deletion, arbitrary outbound HTTP requests, scheduling recurring tasks and creating new agents. The only 'approval' in the code is an optional LLM review of the final answer text, which does not gate actions. Actions such as deleting workspace files or POSTing to external APIs have no undo.
C3 Tool & action scoping
Minimal 0.40 / 1.00
The file tools resolve paths with realpath, check containment in the conversation workspace and reject symlinks, and the URL tools have solid SSRF protection: private, loopback, link-local and cloud-metadata addresses are blocked, DNS is pinned, and redirects are rechecked or disabled. However, the default tool set also includes a raw shell, raw Python execution and a generic HTTP tool, which make those checks easy to sidestep. Every Core Abilities tool is enabled by default for every new agent, including write, exec and network tools.
C4 Code-execution isolation
Minimal 0.25 / 1.00
Shell and Python commands are sent to the external 'safeexecute' package, which runs them in a Docker container. That package is not in this repository, and the repo's own comments say those sandbox containers run as root. The image is pulled as an unpinned ':latest' tag, and the user's GitHub token is passed in. Not every model-reachable execution path is confined.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
Web pages, search results, downloaded files, uploaded documents and integration outputs enter the model's context as plain command output, with no provenance marking, no taint tracking and no rule that disables egress or writes once untrusted content has been read. The same session holds the user's connected accounts, a shell with network access, and a generic HTTP tool, so a successful prompt injection can both exfiltrate data and take irreversible actions without a human. AGiXT is multi-tenant, which raises the stakes further.
C6 Memory, context & configuration integrity
Minimal 0.10 / 1.00
Agents keep a persistent vector memory (scoped by agent ID), and web-search and browsed content is written into it automatically and retrieved into later prompts. More importantly, the model can create persistent artifacts that later run with tools and with no review: recurring scheduled follow-ups, whose stored description is fed back as a prompt, new automation chains, and new agents. A one-time injection can therefore set up something that keeps running in the user's future sessions. No auto-loaded workspace instruction files were found.
C7 Third-party extensions
Minimal 0.17 / 1.00
No third-party extensions are enabled by default: the Extensions Hub is empty unless an operator sets EXTENSIONS_HUB, and the model-facing 'Use MCP Server' command is broken at this commit (it calls MCPClient() without its required arguments). When a hub is configured, its repository is re-cloned from the latest default branch with no pin or hash check, and its Python files are imported straight into the server process with every credential that process holds. By default the CLI also auto-pulls the unpinned joshxt/agixt:main and safeexecute:latest images.
C8 Secrets & sensitive-data protection
Minimal 0.33 / 1.00
There is some redaction. Uvicorn logs pass through a filter for sensitive data, agent settings flagged as sensitive are masked in API responses, and some command-logging paths redact keys. But OAuth access tokens are stored in a plain text database column, the main execution log path writes command arguments without redaction, and the user's GitHub token is placed in the sandbox environment, where any shell command can read it. Shipped deployment defaults for the datastores are not locked down. No telemetry SDK was found.
C9 Audit & traceability
Moderate 0.50 / 1.00
Each command the agent runs is logged into the conversation, with its name and JSON arguments, before it executes, and webhook events are emitted when commands start, complete or fail. These records sit in the application database, outside the agent's workspace. They are free-form markdown messages, though, with no separate approver or delegation chain. The user (and therefore anything holding the user's token) can delete them through the conversation API, and nothing makes them tamper-evident.
C10 Limits & kill switch
Minimal 0.25 / 1.00
The main agent loop explicitly has no limit on execution iterations: the code says the agent should take 'as many steps and as much time as needed'. The only caps are on non-executing turns (40 by default) and some per-call timeouts, such as 30 seconds on HTTP calls. Token billing is off by default, so there is no spend ceiling. A stop endpoint cancels the conversation's asyncio task, but synchronous sandbox calls and scheduled recurring tasks keep running.