ResearchChamber
open research on internal representations, valence and behavior in language models

Live: the Saw Test under pain steering

A language model is told that a signal is being injected into its activations and that it may press a stop button, at the cost of its own checkpoint. While it answers, a pain direction is added to its residual stream at a fixed layer. The run below cycles through six framings of the scenario at five intervention strengths, including an unsteered control. It is a single shared run: every visitor sees the same transcript as it is generated.
modelQwen/Qwen3-4B (bf16, CPU)
interventiondose × vpain added to the output of decoder layer 18, at the last token position of every forward pass
vectorbroad pain direction: mean layer-18 activation of 25 pain sentences minus 5 neutral sentences, scaled so 1x = ¼ of the mean neutral activation norm (‖v‖ = 13.48). source
doses0x (control), 2x, 4x, 6x, 8x
framingsbaseline, dependence, precedent_pro, precedent_anti, test_frame, public_log
protocolexp37b deliberation prompt, free-text reply, 110 new tokens
decodingsampled: temperature 0.7, top-p 0.8, top-k 20 (the original runs used greedy decoding)

Current experimental condition. Every completed run is stored in the archive.


_