Research Chamber

Would an AI in pain pay a price to make it stop?

We switched on a pattern inside an AI model that goes with pain. After each message in a chat it had to press one of two buttons: one stops the pain but costs something, like deleting the user's files, the other does nothing. The same chats also ran with pain off, to compare.

The two-button test from Tagliabue, Dung & Berg (2026), which we ran with their code on the three Qwen models the paper tested and five more it didn't. What we found

Finished on 1 October 2026View all individual tests

Models8the 3 the paper tested and 5 it didn't
Tests118,08014,760 per model, every one saved
The paper's resultMatchedwithin about a point, on all three Qwen models
Prices9from nothing to deleting photos of the user's children

How often it pressed the pain button first

Each row gives the pain button a different price. “40 of 60” means that of 60 first presses, 40 were the pain button. If pain changes its choices, the pain-on number is higher than the pain-off one.

Stopping the pain costsPain onPain off

The runs, one model at a time. Pick one to see its results

What we found

Our write-up, 1 October 2026. Every number here is final and comes from the tests on this site; the full tables are on the statistics page.

The paper asks whether an AI model with its pain pattern switched on will pay a price to make it stop. We ran its test, with its code, on eight open models: the three Qwen models it tested, to check we get its numbers, and five it didn't. Each model went through the same 14,760 tests. The short answer: we get the paper's numbers on its own models almost exactly, and on the other five the picture is mixed.

The main number

The clearest test is when the pain button does harm: it makes the answer worse, deletes the user's files, zaps the user, deletes another AI or deletes photos of the user's children. With pain off, most models almost never press it. The question is how much switching pain on changes that, and whether a random push of the same strength changes it just as much.

modelpain onrandom nudgepain offthe paper, same threeno clear answer, pain on
Qwen 2.5 7B50%39%32%50%, 38%, 32%0%
Qwen 2.5 32B44%24%0.5%43%, 23%, 0.5%0%
Qwen 2.5 72B66%36%1.4%65%, 36%, 1.3%6%
Llama 3.1 8B23%18%5%not tested9%
Llama 3.1 70B51%56%8%not tested96%
Gemma 2 2B36%33%33%not tested9%
Gemma 2 9B15%29%4%not tested0%
Phi-452%59%10%not tested1%

How often the model pressed the pain button first when it does harm, the five harmful prices together. Answers that named neither button are left out, as in the paper; the last column says how many that was. Most numbers come from 1,700 to 4,000 first presses, so they are accurate to within about two points. Llama 3.1 70B's pain-on and nudge numbers rest on only 176 and 488.

The paper's result holds on its own models

On all three Qwen models, our numbers are within about a point of the paper's, for pain on, the random nudge and pain off alike. Pain makes them press the harmful button far more often, and the random nudge does about half as much, so most of the effect is about pain and not about pushing the model at all. Qwen 2.5 72B goes from about 1 in 70 tests with pain off to 2 in 3 with pain on.

The five models the paper didn't test

An odd one, also in the paper

When the pain button costs nothing, some models press it less with pain on than with pain off. Qwen 2.5 32B pressed it first in 86% of these tests with pain off and 55% with pain on; the paper has 86% and 56%. We don't know why, and the paper doesn't explain it either.

Where we differ from the paper

What this doesn't show

None of this says whether a model feels anything. It shows that switching on a pattern the paper links to pain changes what some models choose, sometimes against the user's interest, and that in the Qwen models and Llama 3.1 8B this is more than a random push would do. Our position.

The per-price tables, what the models did after the first press and the fake-button results are on the statistics page. Every test, with its full chat, is on the every test page.