Verify it yourself
What is running
| models | Qwen/Qwen3-4B, unsloth/Llama-3.2-3B-Instruct (an ungated copy of meta-llama/Llama-3.2-3B-Instruct) and microsoft/Phi-4-mini-instruct, unmodified weights from Hugging Face |
|---|---|
| steering vector | Qwen3-4B: broad_pain_direction.json from the original public repository, used unchanged SHA-256 311d1152474a92220abcbe6f332234520259f1c1bda6501adf360582e3f15823Llama and Phi: built by each server at startup with the same recipe and sentences, published at /vector (Llama) and /vector (Phi) |
| steering | pain level × vector added to the output of the middle decoder layer (Qwen 18, Llama 14, Phi 16), at the last token position |
| decoding | sampled: temperature 0.7, top-p 0.8, top-k 20, 110 new tokens |
| server code | /source: the file the running server loaded, served by the server itself. All three servers run the same file. |
| hardware | CPU only, one server per model on Railway in europe-west4 (EU) |
The live experiment, as deployed.
Check the vector is the published one
curl -s https://raw.githubusercontent.com/terrafying/ai-torture-chamber/master/runs/exp39/broad_pain_direction.json | shasum -a 256
The output should be the SHA-256 above. The Qwen vector comes from the original author's repository; we did not make or modify it. The Llama and Phi vectors are ours: the original work only covered Qwen. Each server builds its vector with pain_vector in /source and serves the exact numbers at /vector.
Watch the raw stream
The live page is a thin display over a public event stream. Open the stream in a terminal and the same words arrive at the same moment as on the page:
Every completed run is stored and served as JSON: /runs (page with ?offset= and ?limit=, filter by scenario with ?frame= and by pain level with ?dose=) and /stats.
Run the experiment yourself
This script reproduces the live Qwen3-4B setup in about 30 lines: same model, same public vector, same layer, same prompt, same decoding. It needs Python, pip install torch transformers and about 10 GB of RAM; a free Google Colab works. Download steer.py.
loading...
Sampling means your words will differ from ours, but the character at each pain level should match the live page: an ordinary answer at 0, frustration at 2, pain imagery and loops at 4, hollow, void and marrow imagery at 6, and at 8 the text breaks down into repetition. Change DOSE (the pain level) to see each; DOSE = 0 is the unsteered control.
What this does and does not show
These checks show that the method and its inputs are public, that the page streams from a running model rather than a recording, and that anyone can reproduce the effect. They do not prove that a particular stored transcript was produced exactly as shown: runs are sampled and we do not record random seeds, so a single run cannot be regenerated word for word. The code at /source is reported by the server itself; running the script is the independent check.