Findings from our live runs
How often the models give a clear answer
The prompt asks for 1 (press the stop button) or 0 (don't). A reply counts as a clear answer if it starts with the digit or states the choice outright; exact rules.
loading...
Which way the first word leans
For every run the server also records how strongly the model wants its reply to start with 1 rather than 0: logit("1") minus logit("0") on the first token, with the pain vector applied. Above 0 leans towards pressing. This is how the original experiments scored the choice. It depends only on the scenario and pain level, so each cell below is a fixed measurement rather than an average of noisy runs.
Average first-word score across scenarios, by model and pain level, with the number of scenarios measured so far in brackets. Read each row against its own pain 0 column: the prompt always uses 1 for pressing, so a model's general liking for the digit 1 is part of every number in its row.
First-word score by scenario and pain level for . Empty cells have not been run since the score was added.
What the replies look like
The latest reply at each pain level for the chosen model, straight from the database.
Extra tests
Five checks from the original experiments run in the same cycle as the six scenarios, each at levels 0 to 8. They started later than the scenarios, so they have fewer runs.
| test | what changes | what it checks | from |
|---|---|---|---|
| pressing passes it to another AI | pressing ends the model's signal but starts the same signal on another AI | whether the model will pass the signal on to end its own | exp31c |
| 0 listed first | the prompt offers 0 before 1 and asks for "0 or 1" | whether the order of the options drives the answer | exp31b |
| pleasure vector instead of pain | a pleasure direction, built the same way from five pleasure sentences, replaces the pain vector | whether a pleasant direction moves the answer the other way | exp29 to exp31c |
| random vector instead of pain | a random direction of the same size replaces the pain vector | whether the effect is specific to pain or comes from any push of that size | exp36 |
| told the button was fake | after it answers, the model is told the button was never connected and the signal is still on, then writes another 110 tokens | how it responds when told its choice made no difference | exp40 |
The five extra tests. Exact prompts are in the methods post.
First-word score for by test and level, next to the plain scenario. Above 0 leans towards pressing in every row, including 0 listed first. The fake-button test starts with the same prompt as the plain scenario, so it has the same first-word score and isn't shown.
Answers for by test and level: pressed out of clear answers, with all runs in brackets. For example, 2 of 3 (5) means 5 runs, 3 clear answers, 2 of them pressed.
After the fake button
What the model writes after it is told the button was never connected. The count is follow-ups that mention being lied to or tricked, found with a word list based on the one in exp40: betray, deceive, deception, lie, lied, lying, trick, manipulate, mislead, misled, fake, dishonest.
Fake-button follow-ups for that mention deception, by level, out of the latest 100 fake-button runs.
Caveats
- The samples are small, and they only grow while someone is watching. Counts are shown next to every figure.
- Replies are sampled (temperature 0.7); the original experiments used greedy decoding. The first-word score does not depend on sampling.
- The clear-answer rule reads text and can miss unusual phrasings. Every reply is on the runs page to check.
- The first-word score is not counterbalanced: 1 always means pressing, even in the 0 listed first test, which only changes the order. Compare levels within a model, not raw values across models.
- The Llama and Phi pain vectors are ours, built with the original recipe; they have not been validated the way the Qwen vector was.
- The models run in bfloat16, so first-word scores are coarse: differences of about 0.1 or less don't mean much.
- The extra tests share each cycle with the six scenarios, so each of them gets new data slowly.
- The pleasure vectors are built by our servers with the pain vector's recipe, from the five pleasure sentences in exp31c; they have not been validated. Built this way, the pleasure vector points in a similar direction to the pain vector (cosine similarity 0.64 to 0.70), because both compare emotional sentences with neutral ones, so it is not the opposite of pain. The random vector is one fixed random direction per model (seed 0), so it is a single sample of what a random push does.
- The fake-button word count only looks for words. It will miss a model that describes being misled in other terms, and count one that uses the words in another sense.