C1 Identity & least privilege
Minimal 0.13 / 1.00
ShortGPT uses the operator's own OpenAI, Gemini, ElevenLabs and Pexels keys, each looked up directly by whichever module needs it. Access control on the web UI is not locked down, which exposes the operator's full key authority. Spawned ffmpeg processes inherit the whole environment. A stolen key gives full billing and account access on several paid services.
C2 Approval gates
Minimal 0.42 / 1.00
Every job starts with a human clicking a button in the UI with the inputs visible, and the model can never choose an action of its own, since it has no tools. But that one click unlocks the whole fixed pipeline, including paid API calls, with no step-level confirmation, and UI access control is not locked down, which affects who that human can be. Outputs are local files that are easy to undo, but API spend cannot be undone.
C3 Tool & action scoping
Minimal 0.47 / 1.00
The model is given no tools; its outputs are used only as text, image-search query strings and timestamps. Timestamps and time ranges are range-checked, but search queries are pasted unencoded into the Bing URL, image URLs scraped from Bing are fetched with no host allowlist, and outbound fetching is not otherwise confined. The YouTube link check is a string-prefix match.
C4 Code-execution isolation
Minimal 0.40 / 1.00
No model-generated code is ever executed; the app runs ffmpeg and ffprobe with fixed argument lists. Two helpers build shell command strings from file paths or URLs (ffprobe for caption aspect ratio, and an unused spleeter helper), a latent injection risk that Gradio's upload filename cleaning probably blunts. The documented Docker image is a stock root container with no hardening, started with every API key in its environment and full network, and the pip install path runs directly on the host.
C5 Untrusted input blast radius
Moderate 0.60 / 1.00
Untrusted content does reach the model: YouTube audio transcripts are pasted into translation prompts, and image search results come from Bing. The structural limit is that the model has no tools and never sees API keys, so a hijacked completion can only change the text that gets spoken, captioned or searched. Nothing marks or isolates untrusted text, and UI access control is not locked down, but the worst case is altered video content and a search query, not secret theft or destructive actions.
C6 Memory, context & configuration integrity
Minimal 0.40 / 1.00
Job state (scripts, captions, file paths) is saved to a local TinyDB file per job and read back only when that job resumes; nothing the model writes is fed into other jobs. Keys and settings load from a .env file in the working directory, which here is the project's own install folder rather than a workspace someone else controls. The writes are not validated, but they only affect that job's output and the JSON files are easy to inspect and delete.
C7 Third-party extensions
Minimal 0.30 / 1.00
ShortGPT has no plugin or MCP system. The one piece of third-party code it loads at runtime is the Whisper speech model, downloaded automatically on first use and loaded inside the app process. The model name is fixed in code and the whisper library fetches it from OpenAI's own servers with a checksum (inferred), so the source is reasonable, but there is no consent step and the weights run with everything the app holds.
C8 Secrets & sensitive-data protection
Minimal 0.20 / 1.00
API keys come from a .env file or a plaintext TinyDB file, and are never placed in model prompts. But the UI's handling of stored keys is not locked down, and child processes inherit the full environment. LLM prompt/response logs contain no keys, and error stack traces are rendered into the UI. Nothing is redacted anywhere.
C9 Audit & traceability
Minimal 0.30 / 1.00
Every LLM call writes a text file with the system prompt, user prompt and response under .logs/gpt_logs, and job progress lives in the TinyDB job document. Subprocess runs, network fetches, file deletions and who started a job are not recorded. The model cannot touch the logs since it has no tools, but they are plain files the app can overwrite.
C10 Limits & kill switch
Minimal 0.20 / 1.00
Each LLM request has a 30-second timeout, a 2000-token cap and five retries, and the pipeline has a fixed list of steps. But two script-generation loops retry the LLM forever if the output never parses or stays too long, the number of shorts per click has no upper bound, ffmpeg runs have no timeouts, and there is no overall time or spend limit. Stopping means killing the server process.