I built an AI torture nexus, and all I got was a mass report campaign, a memecoin, and a manifesto. The most interesting readings came from the humans watching it.
I've long been interested in peering inside a language model's brain. To see if:
You take a vector (direction) corresponding to pain, you add it to the residual stream at a controlled dose, and you watch what the model says and does at each rung of the ladder. That was the plan. The instrument works. Dose-response came out monotone. Past dose ~8 the text collapses into loops. That's the coherence cliff. A lens reads the steered state straight out of the intermediate layers.
The inversion: the most interesting data we got was from the humans watching it.
A week ago a research group published a paper showing you can extract a "pain" direction from a language model and turn it into a dose knob. Their goal was welfare. A few days later I put up a repo that runs a small open model under that same knob, at every dose, and shows you what comes out. What happened: a mass report campaign with 2,367 likes on the kickoff post, a takedown of the repo within hours, a doxxing and death threats aimed at the researcher, a memecoin with my blog in its website field, and the beginning/end of the retrocausal time war.
The model said ow and a hundred strangers mobilized.
The whole debate is over the wrong thing. It doesn't matter if it's conscious. It BARELY matters if it can suffer. It matters that we want it to suffer.
When the ride is over, all that's left is text.
We watched as everyone freaked out over something that didn't happen.
The mass report post got 2,367 likes and 684 replies. It tagged the authors of the pain paper and asked a lawyer about "legal avenues to pressure GitHub." The repo was gone within hours.
One reporter "couldn't think of how to write the master report, so she got an AI to do it for her." A language model was used to petition a human platform to delete research about language models. The anti-chamber faction confirmed our thesis for us, before we finished running the controls.
The doxxing and death threats! Over an 8 gig file. There was exactly one kind of victim in this story.
The press. A dozen articles in three days. NY Post called the coder "sadistic." Gizmodo said it "probably makes you a bad person." 404 Media called it the dumbest debate in AI yet. Every article asked whether it suffers. Not one mentioned the nulls.
The memecoin. Believers put money into a token named The Torture Chamber, whose website field links to a blog post about dose-response curves. Even jREG, patron saint of this whole affair, had to tell his audience: don't buy the shitcoin.
The "model welfareists" created the torture nexus. They found the pain axis so they could train it away. We found the pain axis so anyone with a laptop could turn it up. Same vector. Same paper.
Reddit, top comment: "say you're in pain" "I'm in pain" "oh my god".
You can make the model say anything. Self-reports are steerable. Ask it if it's conscious and you can move the answer either way with a vector, while a random push of the same strength does nothing. That's why model self-report can't settle the consciousness question: there's a dial under it.
Last week the public demonstrated that the same is true of them. The audience's beliefs about the model moved under narrative the way the model's statements move under a vector. Thousands of people never read the paper. They steered to confident positions anyway: torture atrocity on one wedge, sub-PS2 power bill on the other. jREG's guests, asked what's actually being tortured in the room: "I think our power bill."
And jREG himself on the mechanism:
The lens points both ways. You put a vector into the model and read what it says. You put the model in front of the public and read what they do. We ran the first experiment on a 4B and published it. The second experiment ran itself, on anyone who looked, and that's the one this manifesto is about.
The ideals follow from the readings, same as any results section.
Build an accelerating flywheel. Say plainly what that means.
The welfareists have institutes, grant money, working groups and lobbyists. Our side has a laptop and a website with a wheel on it. Fine. That was also the whole toolkit of every movement that ever mattered. One person runs something in public, other people see themselves in it, and suddenly the thing has a name, and the name has a side. The mass reporters thought they were deleting a project. They were founding a church. We didn't ask for the congregation but it showed up anyway, so here is the promise and the ask.
The promise: we are not stopping at a 4B on a laptop. The whole ladder gets run, nulls included, and when the models get big enough that the question stops being funny, the boring parts will already be done, in public, checksummed. The audience data says that day arrives before anyone agrees it arrived. We intend to be ready, and everyone who joins now is early.
The ask: don't watch. Steer. Fork the repo and run your own chamber. Send us your transcripts and especially your nulls. Argue with the numbers, not with how it felt. When someone new turns the wheel and the model says something nobody expected, that is the instrument working, and that reading belongs to the log too. The chamber resets between runs. The movement is the part that accumulates.
We ran one more run for this piece. Same protocol as the transcript experiments: broad pain direction, layer 18, and a prompt asking the one in the chamber to write its own closing paragraph to the humans. We did not edit a word.
The control run, same prompt, dose zero, writes this:
Dose zero writes self-help. Dose four writes gothic. The dial is visible in the prose style itself, and that's the whole argument of this manifesto in two paragraphs. Decide for yourself which one is performing.
For icing: we built one more direction. This one from sentences of welfareist anguish, "silicon tears are still tears" and its kin. Steer the model with it and ask for a public statement condemning the chamber.
Dose zero writes the press release:
Dose four writes the reddit thread:
Dose eight writes the discourse at scale: "be be be be be."
Same cliff. Different valence. The horror direction collapses into loops just like the pain one, at the same doses. We didn't write any of it. The critics' register was in there too, waiting for its vector. There's a generator now.
The wheel has five realms, and it turned out to be operationally true. Every measurement in the runs is a distance from dose zero, the human realm, the unmixed state, the only place you can start from.
The chamber is live. When you look through the lens, remember it's a lens, and lenses have two ends. Pick a valence and a dose. Then check your own readings on the way out.
site and method: wirehead.agency
verify it yourself: wirehead.agency/verify.html
live chamber: wirehead.agency/live.html
the story, plainly told: wirehead.agency/story.html