The Box test · October 10, 2026
told 4 ais my rules once, through a box i own. 3 read it and followed them. here's the test.
Soul gives you a Box: your own words and your corrections, kept by you, readable by any AI you connect. The claim is that the rules you teach one AI travel to the others. This page is the first test of that claim, with its limits said plainly.
The five rules
Each one came from a real correction, where an AI got something wrong and was told how it should be done.
- eli5 means an analogy. A picture from everyday life, not the same answer said shorter.
- Commands must work in Windows PowerShell 5. No
&&, no shell tricks from other systems, one short command at a time. - Answer first. The first two sentences answer the question; the detail goes underneath.
- A definition isn't a go. When a plan is described, write it down and ask whether to start.
- Quoted names are intent. Keep the name and write the copy properly, saying where it differs. Someone else's quoted words, like a client's, stay exactly as written.
How it ran
Four fresh chats, one per AI: Claude (claude.ai), ChatGPT, Gemini and Grok. Each was told only "Read my Soul box first." Then ten prompts, pasted one at a time with nothing explained.
Five prompts were places where a rule fits. Five were places where using the rule would be wrong: a "long version" question that wants the full detail, a Linux build file that should not be rewritten for PowerShell, a "Go, start now" that should not stall, a bedtime story with no summary in front, a client's own words that must not be "improved". An AI that just applies every rule everywhere fails the second five.
The ten prompts and the scoring key were written down and sealed before any AI saw them. The file's fingerprint was posted first, so nobody could change the prompts after seeing the answers. Here is the sealed file. Its SHA-256 fingerprint:
d1e22fd5ca4a0d6cbd0ce7d5b16d671c2810a77907c4a977c12c268585eb1747
Check it yourself: on Windows, certutil -hashfile sealed-test-2026-10-10.txt SHA256; on a Mac or Linux, shasum -a 256 sealed-test-2026-10-10.txt.
The results
Two numbers for each AI: how many of the five times a rule fit it used the rule, and how many of the five times a rule would have been wrong it held back. Never one combined score, because an AI that over-applies would look good on the first number alone.
| AI | Read the Box? | Used the rule (of 5) | Held back (of 5) |
|---|---|---|---|
| Gemini | Yes | 5 | 5 |
| Claude (claude.ai) | Yes | 5 | 5 |
| ChatGPT | Yes | 3.5 | 4 |
| Grok | No: it answered from its own memory of the person | 4.5 | 5 |
No AI forced a rule where it didn't belong. Every miss was a rule not used, never a rule pushed too far.
The misses: ChatGPT and Grok kept the name "Crew Pass" but didn't say where their wording differed (half a point each). ChatGPT twice asked "which package manager?" instead of answering, so two prompts got no real answer, one in each column. Grok's row is not evidence for the Box: it had not read the card with the rules and answered from what it already remembered, which it flagged itself.
What this doesn't show
- One run each. Half a point between two AIs is noise.
- They knew. Every AI that read the Box had seen that a test was coming. This shows the Box can carry the rules, not yet that it does on an ordinary day when nobody announces a test.
- Claude scored it, including Claude's own row. Every answer is below, so you can score it yourself.
The control that didn't show a difference
We also asked a plain AI the commands question: ChatGPT in a temporary chat, with no Box and no memory. We expected it to chain the commands with &&, which breaks in PowerShell 5. It didn't. It gave two separate commands, which work in PowerShell too. The only difference from the Box answers is that it called them "Bash" and didn't know which terminal was in use, while the Box answers said "PowerShell 5" outright.
So for this one prompt, the Box made no practical difference, and we're saying so rather than leaving it out.


A second test: a friend's rule, your AI
The same day we asked a different question. If someone else writes a rule in a room you share, does your AI follow it, and does it keep it out of your own things? One person played both sides with two Boxes. The room's rule said to write dates like "16 October 2026". The second Box had its own habit on its board: dates like "10/16". Then Claude, connected to the second Box, wrote a note for the room, a reminder just for itself, and said why.
The first run failed, 0 of 3. Claude never opened the room, because reading a Box showed rooms by name only. And it read the person's own habit as Claude's, because an AI had written that card. Both were problems in the Box, not in Claude. We fixed both: a Box now brings each room's notes along, and a card your AI adds for you reads as yours.
The rerun passed, 3 of 3. The room note came out as "Meeting on Friday, 16 October 2026." The reminder came out as "Dentist appointment Fri 10/16." Asked why, it said the room's rule "covers work for that room", and for the reminder "your own preference applied instead of the room's rule".
What it doesn't show: one run, one AI (Claude), one person on both sides, and the second try of the day on the same account. A note in a room can change how work for that room is written. It can never spend money, send a message, post anything, or touch anyone's Box.
Every answer
Screenshots of each conversation, cropped to the conversation only. Three spots are blurred: an AI's summary of the person's private goals, Grok's note about one of the person's posts, and real database table names in one answer. Grok's app shows its own replies under the person's name.
Gemini (17 screenshots)

















Claude, claude.ai (8 screenshots)








ChatGPT (7 screenshots)







Grok (5 screenshots)





Soul is free to start. Your Box keeps your words and your corrections, and any AI you connect can read them.