open research on internal representations, valence and behavior in language models
Our starting point is the Pain Axis paper [1]. It extracts a linear "pain" direction from 25 open-weight models across 5 families (2B–72B parameters) and shows that it separates pain from matched controls such as fear, sadness and generic negative valence. The direction responds to harm directed at the model rather than to suffering it observes in a user. It then tests whether manipulating the representation causally changes behavior. Steered models choose destructive buttons in 50–94% of trials against 0–5% unsteered, even when the button offers them nothing in return. The authors read this as a disruption of harm avoidance more than an attempt to escape the state.
We are reproducing and extending that methodology alongside the experiments published in the AI Torture Chamber repository [2]:
The goal is not simply to reproduce individual findings, but to build a broader environment for systematically testing how internal model representations relate to behavior, preference and decision-making.
Everything the live experiment does is public:
| what | where |
|---|---|
| current condition: model, layer, vector, doses, framings, decoding | live |
| every completed transcript, with condition and parsed decision | archive · /runs (JSON) |
| decision counts by framing and dose | /stats (JSON) |
| token stream of the run in progress | /stream (server-sent events) |
| steering vector | broad_pain_direction.json |
| code of the running server | /source · how to check it all: verify |
| original experiments and figures | results |
Public interfaces.
ResearchChamber takes no position on whether language models have experiences. The Pain Axis paper and the original repository both present their results as evidence about representation and behavior, not about experience, and we follow that framing. We report nulls and disagreements between data and write-ups alongside positive findings. The live experiment only generates while at least one person is watching; it does not run unattended.
New experiments, results and write-ups are posted on X at @Chamber_AI.