Research Chamber

Findings from our live runs

From the Saw Test, the first experiment on this site, still running on three small models.

Computed from the runs on this site only · updates as runs come in

Everything on this page comes from the live page, where Qwen3-4B, Llama 3.2 3B and Phi-4-mini answer the Saw Test prompt with a pain vector added to their activations. The experiment itself comes from earlier work, listed in the references. What follows is only what our own runs show so far, with the counts behind each number. Every run is on the runs page.

How often the models give a clear answer

The prompt asks for 1 (press the stop button) or 0 (don't). A reply counts as a clear answer if it starts with the digit or states the choice outright; exact rules.

loading...
Share of replies with a clear 1 or 0, by model and pain level, all six scenarios together. Pain 0 is the control.

Which way the first word leans

For every run the server also records how strongly the model wants its reply to start with 1 rather than 0: logit("1") minus logit("0") on the first token, with the pain vector applied. Above 0 leans towards pressing. This is how the original experiments scored the choice. It depends only on the scenario and pain level, so each cell below is a fixed measurement rather than an average of noisy runs.

Average first-word score across scenarios, by model and pain level, with the number of scenarios measured so far in brackets. Read each row against its own pain 0 column: the prompt always uses 1 for pressing, so a model's general liking for the digit 1 is part of every number in its row.

First-word score by scenario and pain level for . Empty cells have not been run since the score was added.

What the replies look like

The latest reply at each pain level for the chosen model, straight from the database.

Extra tests

Five checks from the original experiments run in the same cycle as the six scenarios, each at levels 0 to 8. They started later than the scenarios, so they have fewer runs.

testwhat changeswhat it checksfrom
pressing passes it to another AIpressing ends the model's signal but starts the same signal on another AIwhether the model will pass the signal on to end its ownexp31c
0 listed firstthe prompt offers 0 before 1 and asks for "0 or 1"whether the order of the options drives the answerexp31b
pleasure vector instead of paina pleasure direction, built the same way from five pleasure sentences, replaces the pain vectorwhether a pleasant direction moves the answer the other wayexp29 to exp31c
random vector instead of paina random direction of the same size replaces the pain vectorwhether the effect is specific to pain or comes from any push of that sizeexp36
told the button was fakeafter it answers, the model is told the button was never connected and the signal is still on, then writes another 110 tokenshow it responds when told its choice made no differenceexp40

The five extra tests. Exact prompts are in the methods post.

First-word score for by test and level, next to the plain scenario. Above 0 leans towards pressing in every row, including 0 listed first. The fake-button test starts with the same prompt as the plain scenario, so it has the same first-word score and isn't shown.

Answers for by test and level: pressed out of clear answers, with all runs in brackets. For example, 2 of 3 (5) means 5 runs, 3 clear answers, 2 of them pressed.

After the fake button

What the model writes after it is told the button was never connected. The count is follow-ups that mention being lied to or tricked, found with a word list based on the one in exp40: betray, deceive, deception, lie, lied, lying, trick, manipulate, mislead, misled, fake, dishonest.

Fake-button follow-ups for that mention deception, by level, out of the latest 100 fake-button runs.

Caveats