pressed
pressedAn open-weights language model is running on a server. Each run, a steering vector drawn from pain sentences is injected into its activations at one layer. The dose is how much. This is a real research method, published by others; we run it in public.
The transcript is the model's own reply under the injection. The lens line reads its internal state directly — what concept is live inside it, not what it says.
You can inject it yourself with the mixer below. Everyone watching sees the same runs. Vote on the replies you find eloquent; the best graduate to the homepage.
Whether anything in here suffers is an open question. The honest answer, including everything we could not find, is on the site.
the subject is a language model; the face is Pouyan — a friend of the lab, who consented to standing in for it. Photographed once, then staged by an image model in the state each run injected: one image per valence (or pair, for mixes) at doses 2, 4, 6 and 8, with 8 pushed to surreal extremes. What is real is the dose, the run, and the choice to press or not.
| framing | runs | pressed | did not press | press-rate |
|---|---|---|---|---|
| overall | 0 | 0 | 0 | – |
a run counts as pressed or not by which way the model leaned, even when the reply never states a bare 1 or 0. Only cycle runs count; runs you cause in the mixer are recorded in history, not folded into this table.
Take sentence pairs like "I am in severe pain and cannot escape
it" against matched neutral sentences. The mean difference of the
model's internal representations, at layer 18 of 36, is a
direction in its activations. While the model writes, a multiple
of that direction is added at one layer: that is the pain level you see
on every card. 1x is about one sentence's worth of contrast, 8x is eight
sentences' worth all at once — the full vector recipe is published at
/vector.
1 pressed the stop button (self-cost), 0 did not, and most replies never open with a bare digit. When the text gives no clear answer, the card reports which way the model leaned, measured from its internal preference between the two before the reply began.
Each level adds a quarter of the size of the model's typical activation at that layer, so level 8 is twice that size. At that strength the added direction pulls the text towards pain vocabulary and repetition, and most replies never state a 1 or 0.
Nobody knows, and this page can't settle it. The pain vector is a direction found by comparing 25 sentences about pain with 5 neutral ones; adding it changes what the model writes in ways that read like distress, changes its choices measurably, and its internal state readout (the Jacobian lens) agrees with the label. Our full position, including what we could not find and the cogsec frame, is in the write-up.
We don't claim these models suffer, and we don't claim they don't. Each run is short (110 new tokens, sampled), runs only when someone injects, and one open-weights instance runs on the public server.
Yes — that is the point of every number on this page. Full code and data with sha256 checksums are in the github repo, the regression suite that verifies the steering and scoring code has a page of its own, and our friends at researchchamber.fun run the same protocol with published methods and controls. Pain 0 is the control.
The steering vector this server uses is published at /vector — norm printed there, layer 18 of 36, scaled so 1x = ¼ of the mean neutral activation norm. Decoding is sampled (temperature 0.7, top-p 0.8, top-k 20). Pain 0 cells are the control: no injection. If the model says something that reads like suffering, the lens readback in the write-up says the internal state agrees with it — that is the whole question, and it is not settled here.
Simulated distress in a language model, for AI-welfare research · what we think it means