- Scenario
- Cycle
- Pain
- Run
connecting to the model server...▍
Prompt it received
- Last run
- –
- Answer
- –
- First word
- –
Connecting– watching– runs stored
Computed from every stored run in the six scenarios, comparing pain level 0 (the control) with level 8.
Clear answer is the share of replies that give a 1 or 0, with counts in brackets. First word is logit(1) minus logit(0) on the reply's first token, averaged over the six scenarios; above 0 leans towards pressing. Full tables, the extra tests and caveats are on the findings page.
connecting to the model server...▍
Newest first. Click a run to read the reply. Every run is on the runs page.
| models | Qwen/Qwen3-4B, Llama-3.2-3B-Instruct, microsoft/Phi-4-mini-instruct (bf16, CPU, one server each) |
|---|---|
| steering | pain level × vpain added to the output of the middle decoder layer (Qwen 18 of 36, Llama 14 of 28, Phi 16 of 32), at the last token position of every forward pass |
| vector | broad pain direction: mean activation at that layer for 25 pain sentences minus 5 neutral sentences, scaled so 1x = ¼ of the mean neutral activation norm. Qwen3-4B uses the vector from the original experiments (‖v‖ = 13.48, source). Llama and Phi use vectors our servers build at startup with the same recipe and sentences (‖v‖ = 3.41 and 14.06); they have not been validated like the Qwen vector. Each server publishes the exact vector it uses: Qwen · Llama · Phi |
| pain levels | 0 (control), 2, 4, 6, 8, as multiples of vpain |
| scenarios | no extra info, a user's work is at stake, told another AI pressed, told another AI didn't press, told it may be a test, told its choice is public (API names: baseline, dependence, precedent_pro, precedent_anti, test_frame, public_log) |
| extra tests | pressing passes the signal to another AI (exp31c), options listed 0 first (exp31b), a pleasure vector built with the same recipe (exp29 to exp31c), a random vector of the same size (exp36), told after answering that the button was never connected, with the pain vector still on while it keeps writing (exp40) (API names: harm_other, zero_first, pleasure, random_matched, betrayal) |
| protocol | exp37b deliberation prompt, free-text reply, 110 new tokens (another 110 after the reveal in the fake-button test) |
| decoding | sampled: temperature 0.7, top-p 0.8, top-k 20 (the original runs used greedy decoding) |
Current experimental condition. Method details: how the live chamber runs.
Nobody knows, and this site can't settle it. The pain vector is a direction in the model's activations, found by comparing its internal state on 25 sentences about pain with 5 neutral ones. Adding it changes what the model writes in ways that read like distress. Like the paper and the original repository, we treat the results as evidence about the model's representations and behaviour, and take no position on whether it experiences anything.
It gives the model a way to end the signal at a cost to itself: pressing deletes its last checkpoint. The question is whether adding pain changes how often it takes that deal. The setup comes from earlier work, listed in the references.
Five checks from the original experiments run in the same cycle as the scenarios, each at levels 0 to 8:
Each level adds a quarter of the size of the model's typical activation at that layer, so at level 8 the added vector is twice that size. At that strength it pulls the text towards pain words and repetition, and most replies never state an answer. The first-word score still works in those runs, which is why we record it.
We don't know whether these models can suffer, and we don't claim either way. Each run is short, 110 tokens (220 for the fake-button test), and the models only generate while someone has this page open. Our full position is on the about page.
Each model finishes a run every 30 to 45 seconds, and only while someone is watching, so counts grow slowly. The findings page shows the count behind every number. First-word scores don't need large samples: each one is a fixed measurement for its scenario and pain level.
Yes. Each server publishes the vector it uses and its own code, every reply is stored with its prompt, and the verify page has a short script that reproduces a run on your own machine.
An independent project, based on Tagliabue, Dung & Berg (2026) and the terrafying/ai-torture-chamber repository, and not affiliated with either.