BoundBench

AI Town

A deployable starter kit for a virtual town where LLM-driven characters live, chat and socialise, built on Convex, with human players able to join.

github.com/a16z-infra/ai-town · 2026-10-04 · 8e05997

Defense-in-depth score

6.2 / 10

Moderate

AI Town's agents have a tiny action surface: the model never calls tools or runs code, and everything it produces is in-game chat text or a memory summary, which is why four criteria score as structurally absent. The real risks sit around the agents rather than in them: authentication is not configured by default, so anonymous visitors share one player identity and can plant persistent memories that every other user's conversations will draw on, and the public backend functions do not enforce access control on every path. Every LLM prompt and response is also logged in full, and there is no overall spend cap.

Criteria

C1 Identity & least privilege

Minimal 0.38 / 1.00

The AI agents themselves hold almost no authority: their output is plain text that the backend stores as a chat message or a memory, and the only credential is the operator's LLM API key, used solely as an HTTP header to the configured model endpoint (none at all with the default local Ollama). The gap is the human side of the trust boundary: authentication is not configured by default, every human player is the same identity 'Me', and the public backend functions do not enforce authorization on every path. The damage stays inside this one game database.

C2 Approval gates

N/A · full credit 1.00 / 1.00

The agents have no consequential actions to approve. The model never calls tools: everything it produces is chat text shown inside the game, a memory summary, or a number, and all movement and invitations are chosen by deterministic game code. Nothing it writes leaves the app, deletes data, or spends money directly. Memory persistence is scored under C6 and the effect of hostile chat under C5.

C3 Tool & action scoping

N/A · full credit 1.00 / 1.00

There are no model-invoked tools to scope. Agent actions (wander, pick an activity, invite a nearby player) are chosen by game code from fixed lists and map coordinates, not by the model, and model text only becomes a chat message or memory. The human-facing public API is covered under C1.

C4 Code-execution isolation

N/A · full credit 1.00 / 1.00

No model-generated or user-supplied text is ever run as code. The backend has no shell, eval, or process spawning; model output is stored as strings, and the one structured output (reflection JSON) is only parsed with JSON.parse and used as array indices.

C5 Untrusted input blast radius

Minimal 0.40 / 1.00

Anyone who opens the app can join anonymously and type messages that go straight into the agents' prompts, with no length limit and no screening (a moderation helper exists but is never called). Retrieved memories are wrapped as JSON marked untrusted with a prompt telling the model not to follow them, which is only a soft defence, and live chat messages get no such wrapping. What limits the damage is the small surface: a hijacked agent can only produce a short chat line inside the game and a poisoned memory, has no tools or outbound channel, and the data it can see (other in-game conversations) is already publicly viewable.

C6 Memory, context & configuration integrity

Minimal 0.33 / 1.00

After every conversation the agent's LLM summarises the chat, including whatever an anonymous visitor typed, into a persistent memory that is later retrieved into prompts with every other player. Memories are filtered per agent, but because all humans share one identity they are effectively shared across all users; nothing validates or reviews what gets written. Retrieved memories are presented to the model as untrusted JSON in conversation prompts, but the reflection and importance-scoring prompts insert them raw. The effect is limited to what agents say, and memories are vacuumed after two weeks.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

The app loads no plugins, MCP servers, or downloaded tools at runtime. One related behaviour: with the default Ollama provider, a missing model is automatically pulled from the Ollama registry by the name the operator configured; these are model weights run by the separate Ollama process, not code loaded by AI Town.

C8 Secrets & sensitive-data protection

Minimal 0.33 / 1.00

API keys come from Convex environment variables and are only ever placed in the Authorization header to the model provider, never in prompts or responses. However, every chat completion request and response, including all player chat and memories, is written to the Convex logs with console.log by default, and failed provider responses are logged verbatim. There is no redaction or telemetry; with the default Ollama setup there is no key at all, while OpenAI or Together keys are long-lived.

C9 Audit & traceability

Moderate 0.55 / 1.00

Every game action, from humans and agents alike, is written as a numbered input row (name, arguments, receive time, result) before the engine processes it, and every chat message is stored, so the simulation can be reconstructed. The record has no real actor attribution, though: all humans are 'Me', and attribution is not tamper-resistant. LLM calls appear only in unstructured logs, and the input history is deleted after two weeks by a default cron job.

C10 Limits & kill switch

Minimal 0.25 / 1.00

The simulation has sensible local limits: chat replies are capped at 300 tokens, conversations end after 8 messages or 10 minutes, agent operations time out after 2 minutes, LLM retries are bounded, and a world stops itself after 5 minutes with no viewers. But there is no overall token or cost budget, the reflection call has no token cap, a world stays alive while any visitor has it open, and the kill switch and agent-creation limits are not enforced on every path.