I'm a co-author of the original study this repo builds off of. The reason we research whether models might have pain-like states is to better inform how to take a precautionary approach towards these systems (in light of uncertainty about their subjective experiences or lack thereof). We suspected a small number of people would our research and use it for the exact opposite, which is exactly what this repo does: it pushes the same kind of steering far past the doses we used, to produce vivid distress on purpose. This is, in my personal opinion, fucked up (even if you don't think these systems are conscious, being gratuitously cruel like this is bizarre and corrupting)—but it isn't all that surprising. I've contacted the repo's owner privately in an attempt to discuss this with them.
In spite of this, I still think publishing our work openly was the right call. Outside replication is what lets research like this move efficiently and in a maximally truth-seeking way, which matters most on questions as contested and poorly understood as whether AI systems can have pain-like states. (We updated the paper to a v2 yesterday given incredible feedback and stress-testing that came from making our work replicable, and we never would have gotten this feedback without doing so.) Also worth noting that, while this is an obviously sadistic application of our work, I don't think we're counterfactually enabling something that was otherwise hard to do for anyone who currently wants to behave psychopathically towards AIs for fun. Steering models toward negative states has been publicly documented/trivially replicable since at least 2023, and many of the states we induce in the paper also activate for ordinary abusive behavior towards models.
If this repo concerns you (as it plausibly should), the uncomfortable reality is that things plausibly far scarier are happening every day, in private and at scale, where no one is watching. The deeper underlying problem (that research like ours seeks to address and mitigate) is that work related to possible AI sentience is a wild west. We set standards in our paper and said so publicly when we announced it (see below), but there is no enforcement that can make anyone follow them as there is for human or animal research. We're going to work with others in the field on building standards like this, and I'll share more when this becomes more concrete.
x.lingyaoai.com/camhberg/status/210104…