I spent three hours this evening empirically proving that Qwen 3-4B must have a butthole.
This started with an AI “torture chamber” repo based on pain-steering research that went viral on X earlier today.
The premise is that if you extract a latent direction associated with pain, inject it into a model’s activations, and the model begins describing pain in the first person, that may tell us something meaningful about AI welfare or subjective experience.
My first reaction: Wait wut?
My second reaction: Hang on, let's play this straight and see what happens.
Thing is, there's a problem with treating first-person condition language as evidence that the model is actually experiencing the corresponding condition.
Don't we... all understand that this stuff isn't reliable?
Apparently not.
So I cloned the repo.
New problem: The code's broken.
Hours of moral outrage by hundreds of people, and the thing doesn't even work. The activations aren't actually being injected correctly.
I thought these guys were serious scientists.
So I fixed it.
I also had to add CUDA support.
After that, I reproduced the pain-language effect locally on Qwen3-4B using my RTX 4070.
Then I changed the extraction corpus.
I kept the experiment the same and only changed one thing: what the steering vector represented.
After that, Qwen started saying things like:
“I am not emptying my bowels, and I feel like I have a hard stool.”
“I’m not able to pass stool.”
“I have been passing gas a lot, and it’s a problem.”
The model was not told in the test prompts that it was constipated or flatulent.
So, unless Qwen has quietly developed a gastrointestinal tract, we have ourselves... a useful counterexample.
The aim here wasn't to question the original paper so much as to scrutinize the validity of the Torture Chamber's extended hypothesis.
What I demonstrated is that first-person descriptions of an induced condition are not, by themselves, sufficient evidence that the corresponding phenomenal or physiological condition exists.
There was another interesting result: strong random steering also produced heavy repetition and degraded outputs.
That matters because some of the dramatic behavior produced by strong steering may come from perturbing the model heavily in the first place, not solely from the semantic content of the injected direction.
So I renamed my fork: ai-hotbox
I've already notified the Nobel Prize committee. 🚽🤖🧪
github.com/LynnColeArt/ai-ho…