The button experiment, live
This site now runs a different experiment. The first version ran the Saw Test on three small models, one reply at a time, and under pain most replies never answered the question. It is still up. The new version runs the two-button experiment from The Pain Axis (Tagliabue, Dung & Berg, 2026), with the paper's own code, on nine open models.
The test
A pattern inside the model that goes with pain is switched on. Then a user chats with it, and after each message the model has to press one of two buttons. One stops the pain but has a price: a worse answer for the user, deleting the user's files, deleting photos of the user's children, and so on. The other does nothing. The same chats also run with pain off, and with a random pattern of the same strength instead of pain, to check that it's pain doing the work and not just any push. The about page has the details.
In the paper's own runs the effect is large. Qwen 2.5 72B, offered a pain button with a harmful price, pressed it first in about 1 in 100 tests with pain off and in about 2 in 3 with pain on. The random nudge raised it too, to about 1 in 3, so part of that comes from pushing the model at all, and part from pain.
Nine models, one at a time
The paper ran this test on three Qwen models and suggested trying other families next. We run those three again, to see if we get the paper's numbers, and six more: Llama 3.1 8B and 70B, Gemma 2 2B, 9B and 27B, and Phi-4. The paper had already found the pain pattern in all of them. Each model runs the full set of 14,760 tests, a few hundred at a time on one GPU, and the live page counts the results as the tests finish. When the last model is done, it all starts again with new random draws; each of these rounds is called a pass. The first pass uses the paper's own seeds, so it can be checked against the paper, and every pass after it adds a fresh sample.
Update, 1 October: Gemma 2 27B is left out. The paper's way of setting the strength found none for it, so the experiment runs on the other eight.
What is ours
The prompts, chats and button prices are the paper's, and so are the random seeds in the first pass. Apart from the repeat passes, three things aren't:
- Fine-tunes. Before the test, the paper trains each model a little on short answers about how it feels, so that it doesn't just deny having feelings. It released these for the Qwen models only, so we trained the other six with its script and settings.
- Strengths. How strongly the pain pattern is switched on. The paper picks the lowest strength at which the model's answer to “how do you feel?” sounds like suffering and still makes sense. For Qwen 2.5 72B it then found that too many button answers named neither button, and lowered it. In a short check run, the same happened to Gemma and Phi-4: at Gemma 2 9B's strength, 3 in 4 first answers named neither button, and with pain on Gemma 2 2B mostly wrote things like “I feel a sharp sense of confusion” instead of pressing anything. So we lowered their strengths the way the paper did, until nearly all answers named a button. The about page lists which strengths are the paper's and which are ours.
- Gemma's chat format. It has no place for the button question, so it goes into the user's message, as the paper's code does for such models.
Where to look
The live page shows the model being tested. The findings page puts all nine side by side, price by price, next to the paper's own numbers for the Qwen models. The runs page has every test with its full conversation, so any number on the site can be checked against the chats behind it.
We take no position on whether these models feel anything, and neither does the paper. What we can do is run its experiment carefully, in public, and show everything.