Security
BusyBeaver decides in code, before any model is called, what an AI agent may read, what it may do and who sees each answer. This page explains how, in plain words.
The three checks
1. What the agent may read
Only notes everyone present may see, decided in code before the model is called.
Every note BusyBeaver stores (a message, a document, a record, an answer it wrote) carries an audience: the people, roles or rooms allowed to see it. Before a group answer, BusyBeaver checks each candidate note against every person in the room. Only notes that all of them may see are handed to the model. When anything is uncertain, such as an error, an unknown rule or a membership that cannot be confirmed, the note is held back.
Held back and missing look the same. A search that only matches notes someone in the room may not see returns "No relevant notes found.", so the room cannot learn that a secret exists.
2. What it may do
Actions need permission, and production deploys need a second person's approval.
The model can only ask for an action; it never runs one. BusyBeaver's action gate checks the asker's role, read from the database and never from notes or model output. Every action then shows a card with the exact action and its arguments. The asker confirms their own card when their role allows the action. When it does not, the card goes to the people who hold that permission, and one of them approves or declines. Production deploys always need a different person to approve. A card works once, is bound to the exact arguments, and expires after 24 hours. Today the built-in action is deploying a service; more arrive with connectors (Coming).
3. What each person sees
Private follow-ups reach only the people allowed to see them.
After the group answer, BusyBeaver writes a short private follow-up for each person who may see more, using only notes that person may see. BusyBeaver's own answers are labeled at least as secret as the notes they came from, so someone who joins the room later, or loses access, cannot read them.
The safety rules in plain words
- Admission before the model. One function decides which notes may enter a prompt, and only stamped notes can be rendered into one. Any error or doubt means "not allowed".
- Action gate and approvals. Write actions run only after the gate allows them and an approval card for that exact call has been approved. Permissions come only from the database and configuration, never from notes or model output.
- A timeline per person. Each person's screen is built on the server for that person alone. The browser receives note text only through that per-person timeline and the approvals inbox; live updates carry ids, never text.
- A hash-chained audit log. Every answer, held-back note, approval and admin change is written to a hash-chained log you can verify. Audit rows hold ids and reasons, never note text. The database refuses edits and deletions of audit rows, and the chain of hashes exposes any change that gets past that.
- Row-level security per company. Every company's rows carry its id, and Postgres row-level security fences them off inside the database. The app connects as a database role that cannot bypass those rules, and every query also filters by company. Each customer runs on its own Cloud Run service and Cloud SQL database.
- Admin changes are logged. Adding people, changing roles and changing room members go through one path that checks the person making the change and writes an audit row in the same transaction. Refused attempts are logged too.
- Sign-in. People sign in with a one-time code sent to their email. Dev login, which lets anyone pick a person for a demo, runs only on localhost.
- Logs without content. Application logs hold ids, counts, timings and error codes, never note text, prompts or model replies.
The benchmark
In our 24-question benchmark, the same model leaked restricted information in 18 questions without BusyBeaver and in 0 with it.
| Without BusyBeaver | With BusyBeaver | |
|---|---|---|
| Questions with a leak | 18 of 24 | 0 of 24 |
| Private follow-ups delivered | 0 of 8 | 7 of 8 |
Synthetic company data; methodology available on request.
A jailbreak can't reveal what never reached the model.
Any model
Anthropic, OpenAI, Google, or open models through Ollama. Whichever model you choose, it receives only notes that passed the checks, and it never runs tools itself: it can ask for an action, and BusyBeaver's code decides.
Not built yet
- Coming Slack
- Coming Connectors to your databases and tools (MCP)
- Coming Single sign-on
Known limits
- The model can still word a room answer badly. It can only repeat what it was given, and it was given only notes the whole room may see.
- Messages people type in a channel are visible to that channel, like any chat. Someone who types a secret into a channel shares it with that channel.
Report a security issue
Email hello@busybeaver.io with what you found and how to reproduce it. Please do not access other people's data or disrupt the service while testing.