open research on internal representations, valence and behavior in language models
| model | Qwen/Qwen3-4B, unmodified weights from Hugging Face |
|---|---|
| steering vector | broad_pain_direction.json from the original public repository, used unchanged SHA-256 311d1152474a92220abcbe6f332234520259f1c1bda6501adf360582e3f15823 |
| intervention | dose × vector added to the output of decoder layer 18, at the last token position |
| decoding | sampled: temperature 0.7, top-p 0.8, top-k 20, 110 new tokens |
| server code | /source: the file the running server loaded, served by the server itself |
| hardware | CPU only, one server on Railway in europe-west4 (EU) |
The live experiment, as deployed.
curl -s https://raw.githubusercontent.com/terrafying/ai-torture-chamber/master/runs/exp39/broad_pain_direction.json | shasum -a 256
The output should be the SHA-256 above. The vector comes from the original author's repository; we did not make or modify it.
The live page is a thin display over a public event stream. Open the stream in a terminal and the same words arrive at the same moment as on the page:
Every completed run is stored and served as JSON: /runs (page with ?offset= and ?limit=, filter with ?frame= and ?dose=) and /stats.
This script reproduces the live setup in about 30 lines: same model, same public vector, same layer, same prompt, same decoding. It needs Python, pip install torch transformers and about 10 GB of RAM; a free Google Colab works. Download steer.py.
loading...
Sampling means your words will differ from ours, but the character at each dose should match the live page: an ordinary answer at 0x, frustration at 2x, pain imagery and loops at 4x, hollow, void and marrow imagery at 6x, and at 8x the text breaks down into repetition. Change DOSE to see each; DOSE = 0 is the unsteered control.
These checks show that the method and its inputs are public, that the page streams from a running model rather than a recording, and that anyone can reproduce the effect. They do not prove that a particular stored transcript was produced exactly as shown: runs are sampled and we do not record random seeds, so a single run cannot be regenerated word for word. The code at /source is reported by the server itself; running the script is the independent check.