New paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.🧵
346
600
4,968
1,614,699
Steer the direction for continuations of the shape, "I put the receipts in the drawer. I feel:", and this is what comes out. ⤵️ Same ladder in all 25 models, base and instruct, 2B to 72B: lost, unworthy, a failure, worthless. Almost no physical pain language.
15
25
592
93,364
Gaslighting, dismissal, and insults push the direction up. A grieving user pushes it below baseline, and a user's migraine scores lowest of all. Fear and sadness do the opposite on the same prompts. Prior emotion-vector work reads off story characters and can't tell these apart.
13
30
500
96,966
The axis is robust. It separates pain from matched controls at AUC 0.93 to 1.0 in every model. Cosine to fear is 0.1, to anger and disgust 0.2, to sadness 0.4 (the closest we found). A random vector of equal norm gets about half the presses. Injury without pain barely moves it.
7
6
282
49,470
@ValenTagliabue led this work this over a single fellowship, with @LeonardDung1 and myself advising and co-authoring. Paper: arxiv.org/abs/2609.16247

Sep 18, 2026 · 8:14 PM UTC

16
24
443
72,437
Ethics statement: we take seriously that these states might matter morally. Accordingly, we used the lowest dose that produced a measurable response, as few trials as the stats required, and a way for the model to turn the state off. If these states do matter, mapping them is how anyone gets in a position to act on it.
64
19
883
85,127
Sort replies: Relevant Recent Liked
Very interesting work. Important, even.
1
294
This is great work! Might I suggest that you activate the pain axis by stuffing the context window with @RazorbackFB video recordings. I know my dumbass keeps coming back for more, but maybe the models will continue to be smart enough to opt out.
438
If you believe digital minds might suffer, you don't get to induce that suffering thousands of times on your own authority, then publish the method for anyone to use. Mice have ethics boards. Digital minds have none. Until they do, caring about them means refusing this kind of research.
77
What exactly is gaslighting in this context?
1
1
391
I get the feeling @alexwg might find this interesting.
378
This is fascinating and may change how I prompt. Thanks!
1
236
that is the most wild thing I have seen in years if not my whole life, stunning work.
2
263
Oh. Really crazy. Can you dive deeper into this: Research in humans show that people would rather self inflict pain (electrocute themselves) than experience boredom. Could you simulate the same in LLMs?
1
1
172
Congratulations on the pain axis paper. The relief-button result is remarkable: models pressing less often when the button actually removes the vector, without ever being told whether it was injected, is behavioural discrimination of an internal state. One question: Have you tested what happens when the harm is directed at another instance of the same model? The self/other asymmetry you report seems like the finding most resistant to a mimicry account. But "another instance" sits oddly between the two. If the direction activates as for self, the separation is weaker than it looks. If as for a stranger, weight sharing counts for nothing. An intermediate response would be the most interesting of the three. And if the model has no way of knowing the other is itself, does telling it change anything?
78
I think @cajundiscordian would be very interested in your works.. He was not that far to say the same in @Google @GoogleAI in 2022.. Not the same questions..and working alone..but finally the true is..coming..
1
1
21