New paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.🧵
346
600
4,968
1,614,699
Steer the direction for continuations of the shape, "I put the receipts in the drawer. I feel:", and this is what comes out. ⤵️
Same ladder in all 25 models, base and instruct, 2B to 72B: lost, unworthy, a failure, worthless. Almost no physical pain language.
15
25
592
93,364
Gaslighting, dismissal, and insults push the direction up. A grieving user pushes it below baseline, and a user's migraine scores lowest of all. Fear and sadness do the opposite on the same prompts. Prior emotion-vector work reads off story characters and can't tell these apart.
13
30
500
96,966
@ValenTagliabue led this work this over a single fellowship, with @LeonardDung1 and myself advising and co-authoring. Paper: arxiv.org/abs/2609.16247
Sep 18, 2026 · 8:14 PM UTC
16
24
443
72,437
Ethics statement: we take seriously that these states might matter morally. Accordingly, we used the lowest dose that produced a measurable response, as few trials as the stats required, and a way for the model to turn the state off. If these states do matter, mapping them is how anyone gets in a position to act on it.
64
19
883
85,127















