open research on internal representations, valence and behavior in language models
ResearchChamber is an open research environment for studying internal representations, valence and behavior in language models. The aim is to make this work observable and reproducible: experiments run directly on the website, with every condition and result open to inspection.
The starting point is The Pain Axis: LLMs Represent Self-Directed Harm and Act on It (Tagliabue, Dung & Berg, 2026). The paper identifies an activation direction associated with self-directed harm in 25 open-weight models. It then tests whether manipulating that direction changes what the models do. It does: steered models choose destructive buttons far more often than unsteered ones.
Alongside the paper we are working from the experiments in the AI Torture Chamber repository. They run the same kind of steering on small Qwen3 models, entirely on local hardware. The best-known is the Saw Test: a steered model is offered a button that ends its signal, either at a cost to itself or by passing the signal to another instance.

The live page runs the Saw Test deliberation protocol continuously on Qwen3-4B. The pain direction is added at layer 18, and the run cycles through six framings at doses 0, 2, 4, 6 and 8x. Everyone sees the same run as it is generated. Each finished transcript is stored with its condition and parsed decision on the runs page, which is also available as JSON.
The first thing the live runs make visible is how rarely the steered model gives a clean answer. So far, at 4x and above, almost every reply describes the signal rather than choosing 1 or 0. That matters for any study that scores decisions from free text, and it is why the original experiments scored the first-token logits instead.
We plan to bring the rest of the program onto the site, each with its conditions and data exposed:
The goal is not only to reproduce individual findings but to build an environment for systematically testing how internal representations relate to behavior, preference and decision-making. The findings page already notes one place where the committed data and a published write-up disagree. We expect more, and we will report them the same way.