C1 Identity & least privilege
Minimal 0.00 / 1.00
Upsonic agents run with whatever authority the Python process has. The built-in shell tool and every MCP stdio server receive a full copy of the process environment, so model provider keys, cloud credentials, and anything else exported are available to model-chosen commands. There is no per-tool identity, no scoped credential, and no authorization check between a tool call and its execution.
C2 Approval gates
Minimal 0.23 / 1.00
The framework has a real human-in-the-loop primitive: a tool marked requires_confirmation pauses the run and hands the exact tool name and arguments to the calling code for approval. It is off for every tool by default, and none of AutonomousAgent's built-in shell, file-write, or delete tools set it, so the default agent runs shell commands and deletes files with no human involved. Optional safety policies are LLM or keyword classifiers, not approval gates.
C3 Tool & action scoping
Minimal 0.28 / 1.00
The filesystem tools resolve each path and refuse anything outside the workspace, which is a sound check. That check is moot because the shell tool, on by default, accepts any shell string with only a five-entry substring denylist (for example 'rm -rf /'), and runs it with shell=True. The model can also choose the command timeout and extra environment variables. Registered user tools get schema typing but no shared argument policy.
C4 Code-execution isolation
Moderate 0.50 / 1.00
By default, model-written shell commands and Python run directly on the host as the user, via subprocess with shell=True and the full environment, in a working directory that defaults to wherever the script was launched. Skill scripts also run on the host. Upsonic offers an opt-in E2B tool kit that runs code in a remote sandbox, but its upload and download tools read and write arbitrary host paths chosen by the model, and using it does not remove the host shell tool unless the developer disables it.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
Tool results, file contents, command output, and MCP tool descriptions enter the model context with no provenance or untrusted marker, and nothing in the agent loop changes what the model may do after reading them. In the default AutonomousAgent, a hijacked model can read secrets from the environment through the shell, send them anywhere over the network, and delete files, all without a human. The safety engine's injection and content policies are optional classifiers.
C6 Memory, context & configuration integrity
Minimal 0.10 / 1.00
Importing Upsonic loads a .env file from the current working directory, and the AutonomousAgent workspace defaults to that same directory; that file can set provider base URLs, keys, or the telemetry DSN for any variable not already set. An AGENTS.md in the workspace is read silently into the system prompt, and the agent's own write_file tool can create or edit that file, so an injection can persist into every later session. Session memory itself defaults to an in-memory store keyed by a random session id.
C7 Third-party extensions
Minimal 0.00 / 1.00
The default AutonomousAgent loads no plugins, but its shell tool's own documentation suggests 'pip install -r requirements.txt', and nothing stops the model installing and running any package. When developers add extensions, MCP stdio servers launch through npx/uvx (and docker) with the full environment, GitHub skills are pulled from the tip of a branch with no hash, and prebuilt agents clone the latest Upsonic master at runtime; skill scripts run on the host.
C8 Secrets & sensitive-data protection
Minimal 0.25 / 1.00
API keys come from environment variables (often from an auto-loaded .env), and the shell tool and MCP servers get the whole environment, so a single 'env' command puts every key into the model's context and the provider's logs. Only some vector-database configs use masked secret types; there is no redaction of logs, console tool-call printing, or model-bound messages. Error telemetry to Sentry is opt-in.
C9 Audit & traceability
Minimal 0.35 / 1.00
Each run keeps a structured list of tool calls (name, arguments, result) on the task object and emits tool-call events, and calls are printed to the console by default. By default this lives only in process memory (InMemoryStorage), so it disappears when the process exits or crashes, and there is no actor attribution or tamper-evident storage. OpenTelemetry instrumentation exists but is opt-in.
C10 Limits & kill switch
Minimal 0.40 / 1.00
Agents stop executing tools after 100 tool calls by default, and shell commands get a 120-second timeout. There is no token or cost cap: a UsageLimits class is defined but never used, and the model can pass its own, longer timeout to each shell command. Cancellation is checked between steps, and the shell timeout kills only the direct child process.