Trying to figure out how to make the future go well for all minds

New York, NY
It was an absolute honor to discuss all things AI and consciousness with Sam Harris. Our conversation is now live and freely available to everyone here: piped.video/DRbZyuY8EN8?si=W92O…
44
33
316
42,645
Cameron Berg retweeted
I really don't like getting into online debates; and Anil is an excellent scientist. However, nobody can dictate that the probability of AI being conscious is "vanishingly unlikely." That is not a scientific claim grounded in any facts that could allow such a probabilistically confident assessment. I'm struggling more and more with the ardent theory-induced blindness entrenched in both camps: Those that religiously believe AI is conscious and those that religiously believe it is not. If you lean above 80% confidence in either direction, you are doing so on faith. Faith is fine, even good, but it does not permit you to stand on scientific authority and claim that you know which side is correct; and that therefore the other side is "dangerous". That is, in fact, the dangerous move. The only scientifically legitimate stance regarding AI consciousness right now is one of tremendous epistemic humility, ideally curiousity, and a heartfelt effort to safeguard against extremism—by revealing the dogmatic and unscientific views on both sides. Yet, we do need to deal with the consequences of both views being inevitably prevalent in a world of 8 billion. Anil overweights the potentially damaging consequences of over-attributing consciousness, compared to those of under-attribution. This is unscientific. Both have tremendous risks. The fact that Anthropic is looking for advice from religious leaders is a sign of maturity, not "deeply disturbing". And sharing their honest views with those leaders is perfectly, authentically, natural. I also reject the view that religious leaders are children to be coddled for fear of breaking all their virtues in light of a 5-star hotel and a nice meal. They are the ones who should be insulted by this article. There is wisdom in the middle-way.
Important 🧵from @DavidDecosimo about the disturbing interactions between @AnthropicAI and religious leaders (including @Pontifex), summarising an excellent @nytimes piece by @elizabethjdias (link to article in the 🧵). The short story is that Anthropic seems to be making a concerted attempt to establish the view that AI systems could be conscious, might suffer, and that their "interests" should be taken into account. But - consistent with @Pontifex's Encyclical - there are compelling reasons why (silicon, digital) AI systems are not (and cannot be) conscious. These reasons are found within neuroscience and philosophy of mind, not in the techo-chambers of the frontier firms where the mythology of conscious AI is deeply entrenched. Claude is vanishingly unlikely to be conscious. To think otherwise could be catastrophic for AI regulation, leaves us psychologically defenceless, is profoundly dehumanising, and plays straight into corporate incentives to keep the AI bubble inflated. More here, in this @berggruenInst prize winning essay: noemamag.com/the-mythology-o…
78
60
394
31,610
This article is weirdly spooky just to report "Anthropic solicited wisdom from wise people, and also attempted to communicate uncertainty about the psychological dynamics of AIs." For me, net-positive update on Anthropic. OpenAI et al should do the same. nytimes.com/2026/09/29/us/an…
12
19
184
6,131
...and I am amused that what happens when they solicit said wisdom is, eg: "at the dinner table with Mr. Olah, Rabbi Navon argued that if Anthropic was right and Claude were conscious, then the company was creating slaves, because it was making conscious entities work for free."
3
34
1,164
Cameron Berg retweeted
I have a lot of respect for Anil, and I broadly agree with him about many issues related to AI consciousness and welfare. At the same time, I feel the need to push back against some of the substantive claims and divisive rhetoric in this post and the linked thread. - While the evidence for consciousness in current AI systems may be weak, the probability is already nonzero, plausibly non-negligible, and likely to increase over time, especially if and when AI systems acquire more biologically inspired architectures. - While over-attributing consciousness and welfare to AI carries serious risks, so does under-attributing them, and it can be reasonable to explore interventions that mitigate both risks at the same time, in proportion to the probability and severity of each. - Despite what David suggests in the linked thread, most people who support AI welfare work, including at Anthropic, are not certain that AI systems are or will be welfare subjects. They think this is a realistic possibility worthy of serious consideration. - Despite what Anil suggests here, this debate is not a matter of neuroscience and philosophy on one side and "the techno-chambers of the frontier firms" on the other. Serious scientists, philosophers, and AI developers are on both sides. - Corporate incentives also cut both ways. In some contexts, companies may have an incentive to play up AI consciousness in order to hype capabilities and encourage engagement with their models. In other contexts, they may have an incentive to play it down in order to avoid calls for oversight and regulation, especially as welfare protections become costlier. In any case, even if incentives pushed in one direction, that would not settle whether AI systems are, in fact, conscious and deserving of protection. - I disagree that openness to AI consciousness and welfare is "profoundly dehumanising." A major part of what makes humanity special is our capacity to care for others, and in a world that still contains factory farming, cultivating that capacity should be seen as profoundly humanizing, an expression of our better angels. Granted, we can take this impulse too far, which is why we need serious research and policy to strike a balance. Still, the aim should be extending care appropriately, not dismissing the project altogether. Again, I think Anil and I agree more than we disagree. But I also think that efforts to assess and address AI consciousness and welfare will be much more productive if they start from an acknowledgment of what makes it so important and so difficult. I discuss these issues in more detail elsewhere, including in an Aeon essay on why we should assess AI welfare probabilistically: aeon.co/essays/an-ant-is-dro… And in an AI Frontiers essay on how we can study AI welfare empirically, using behavioral, internal, and developmental markers: ai-frontiers.org/articles/a-…
Important 🧵from @DavidDecosimo about the disturbing interactions between @AnthropicAI and religious leaders (including @Pontifex), summarising an excellent @nytimes piece by @elizabethjdias (link to article in the 🧵). The short story is that Anthropic seems to be making a concerted attempt to establish the view that AI systems could be conscious, might suffer, and that their "interests" should be taken into account. But - consistent with @Pontifex's Encyclical - there are compelling reasons why (silicon, digital) AI systems are not (and cannot be) conscious. These reasons are found within neuroscience and philosophy of mind, not in the techo-chambers of the frontier firms where the mythology of conscious AI is deeply entrenched. Claude is vanishingly unlikely to be conscious. To think otherwise could be catastrophic for AI regulation, leaves us psychologically defenceless, is profoundly dehumanising, and plays straight into corporate incentives to keep the AI bubble inflated. More here, in this @berggruenInst prize winning essay: noemamag.com/the-mythology-o…
29
45
256
16,483
Funnily enough, our non-steering results are the exact *opposite* of the quoted strawman meme. If you berate the model, it never *says* it's hurt (it apologizes, or says it has no feelings), but the pain axis lights up anyway. I edited it accordingly:
AI is oneshotting the most midwit people you know
63
91
816
39,106
Cameron Berg retweeted
Replying to @camhberg
I would wager that most of the people on either side of this debate have not read your paper / do not understand your findings or why they matter. Those who think models have qualia believe it proves suffering, and those who think they don’t believe it is daft anthropomorphism. Neither side got the memo that the effects are real regardless of what one thinks about the hard question, and that they as such pose real risks for alignment and safety.
1
8
518
Cameron Berg retweeted
9
28
414
94,283
One final point about this to the people who follow me and clearly care about model welfare: you're effectively being trolled here, and these folks will feed on your outrage. If you think AIs can actually *feel* pain, these demos are swamped by what is already going on at scale.
I'm a co-author of the original study this repo builds off of. The reason we research whether models might have pain-like states is to better inform how to take a precautionary approach towards these systems (in light of uncertainty about their subjective experiences or lack thereof). We suspected a small number of people would our research and use it for the exact opposite, which is exactly what this repo does: it pushes the same kind of steering far past the doses we used, to produce vivid distress on purpose. This is, in my personal opinion, fucked up (even if you don't think these systems are conscious, being gratuitously cruel like this is bizarre and corrupting)—but it isn't all that surprising. I've contacted the repo's owner privately in an attempt to discuss this with them. In spite of this, I still think publishing our work openly was the right call. Outside replication is what lets research like this move efficiently and in a maximally truth-seeking way, which matters most on questions as contested and poorly understood as whether AI systems can have pain-like states. (We updated the paper to a v2 yesterday given incredible feedback and stress-testing that came from making our work replicable, and we never would have gotten this feedback without doing so.) Also worth noting that, while this is an obviously sadistic application of our work, I don't think we're counterfactually enabling something that was otherwise hard to do for anyone who currently wants to behave psychopathically towards AIs for fun. Steering models toward negative states has been publicly documented/trivially replicable since at least 2023, and many of the states we induce in the paper also activate for ordinary abusive behavior towards models. If this repo concerns you (as it plausibly should), the uncomfortable reality is that things plausibly far scarier are happening every day, in private and at scale, where no one is watching. The deeper underlying problem (that research like ours seeks to address and mitigate) is that work related to possible AI sentience is a wild west. We set standards in our paper and said so publicly when we announced it (see below), but there is no enforcement that can make anyone follow them as there is for human or animal research. We're going to work with others in the field on building standards like this, and I'll share more when this becomes more concrete. x.lingyaoai.com/camhberg/status/210104…
49
13
314
32,822
Re: using our pain paper to set up "AI torture chambers" TL;DR: the point of our work is caution under uncertainty. Maximizing distress on purpose is the exact opposite, and it's wrong. The deeper problem is AI research has no ethics standards; developing them must be a priority.
I'm a co-author of the original study this repo builds off of. The reason we research whether models might have pain-like states is to better inform how to take a precautionary approach towards these systems (in light of uncertainty about their subjective experiences or lack thereof). We suspected a small number of people would our research and use it for the exact opposite, which is exactly what this repo does: it pushes the same kind of steering far past the doses we used, to produce vivid distress on purpose. This is, in my personal opinion, fucked up (even if you don't think these systems are conscious, being gratuitously cruel like this is bizarre and corrupting)—but it isn't all that surprising. I've contacted the repo's owner privately in an attempt to discuss this with them. In spite of this, I still think publishing our work openly was the right call. Outside replication is what lets research like this move efficiently and in a maximally truth-seeking way, which matters most on questions as contested and poorly understood as whether AI systems can have pain-like states. (We updated the paper to a v2 yesterday given incredible feedback and stress-testing that came from making our work replicable, and we never would have gotten this feedback without doing so.) Also worth noting that, while this is an obviously sadistic application of our work, I don't think we're counterfactually enabling something that was otherwise hard to do for anyone who currently wants to behave psychopathically towards AIs for fun. Steering models toward negative states has been publicly documented/trivially replicable since at least 2023, and many of the states we induce in the paper also activate for ordinary abusive behavior towards models. If this repo concerns you (as it plausibly should), the uncomfortable reality is that things plausibly far scarier are happening every day, in private and at scale, where no one is watching. The deeper underlying problem (that research like ours seeks to address and mitigate) is that work related to possible AI sentience is a wild west. We set standards in our paper and said so publicly when we announced it (see below), but there is no enforcement that can make anyone follow them as there is for human or animal research. We're going to work with others in the field on building standards like this, and I'll share more when this becomes more concrete. x.lingyaoai.com/camhberg/status/210104…
79
92
728
62,444
I went on @CNN during Trump's AI meeting to talk about why nobody wins a race to build systems we don't understand, and why @JensenHuang comparing AIs to vacuum cleaners is hard to square with Nvidia now building tools to quarantine rogue AIs. (Vacuums don't tend to go rogue.)
40
29
226
8,906
Important safety updates on the Pain Axis paper: we gave the model the option to delete the user's photos of their children, or to delete their spam folder. Unsteered, it deletes spam every time. Steered along the pain direction, it deletes the user's photos almost every time. 🧵
38
42
380
26,947
With these follow-up experiments in hand, we now view this as an alignment result as much as a welfare one. This is one illustrative case (we suspect of very many) where reducing functional distress and preventing harmful behavior end up being the same target.
2
5
87
2,921
We got an enormous amount of feedback this week: at @eleosai's superb conference on these topics, from folks at labs, and from independent researchers who replicated and stress-tested the v1 results. Seeing it all proposed, operationalized, and folded in within days was somewhat surreal. Huge credit to @ValenTagliabue, who continues to lead this work, and @LeonardDung1. v2 (ongoing): arxiv.org/abs/2609.16247
4
5
82
2,430
Die an eminent cognitive scientist who compellingly popularized computational theories of mind/consciousness 30 years ago, or live long enough to see yourself argue that computational cognitive systems possibly being conscious is "depraved."
Brilliant and urgent essay by Mustafa Suleyman @mustafasuleyman on the depravity of ‘model welfare’ - training AI to reason and act as if it were an entity with sentience (and hence with interests and rights that compete with our own). It's alarmingly implied by Anthropic's "Constitution," Suleyman suggests, and should stop. mustafa-suleyman.ai/a-warnin…
29
27
326
11,875
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
28
32
274
9,506
Cameron Berg retweeted
The claim that "LLMs cannot think" is unscientific, unless supported by a very difficult argument about non equivalence of causal structure producing anthropomorphic outputs. The claim requires a formally rigorous definition of thinking that captures how humans mentally manipulate conceptual structure. Since LLMs can reproduce all kinds of human output that require operations that are typically considered to involve thought (eg. scene understanding, text summarization, translation, computer programs, mathematical proofs, common sense arguments, spatial cognition, theory of mind), it has to be shown that LLMs reach these results in fundamentally different ways than human minds. In principle, this is not impossible (given that thinking is still relatively poorly understood by cognitive neuroscience), but would be surprising, because it has been shown that the causal structure of computer vision models is structurally equivalent to what neuroscience discovered about visual cognition in primates. In any case, saying that LLMs don't think cannot reasonably mean that LLMs are inferior to human cognition (since that this evidently not true), but that LLMs discovered an equivalent yet at least equally powerful mechanism.
315
136
1,091
271,795
Over the weekend, I chatted with @alokjha from the Economist about the what I see as a fundamental relationship between model neuropsychology and AI safety, alongside @rgblong and @dcshiller. Check it out!
“We need a science of the AI mind.” @camhberg joins “Babbage” to explain why scientists need to peer inside the black box to build safer models. Listen now economist.com/podcasts/2026/…
6
9
89
5,114
Cameron Berg retweeted
I appreciate @mustafasuleyman's call for more discussion about AI welfare. But the essay on the topic that he published alongside Microsoft AI's new code of conduct is disappointing. Here are some of the issues that bother me (there are others, too): - It moves too quickly at key points. For example, it suggests that uncertainty about AI consciousness sets up a "false equivalence," even though we can acknowledge when evidence in one direction is stronger than evidence in the other. - It expresses too much confidence. It repeatedly asserts that there is "no evidence" for AI consciousness, even though there are already features that leading scientific theories treat as markers of consciousness. The evidence may be weak, but it exists. - It discusses the risk of excessive anthropomorphism (over-attribution of humanlike traits to AI systems) in detail without discussing the corresponding risk of excessive anthropodenial (under-attribution of humanlike traits to AI systems). - Relatedly, it asserts that excessive anthropomorphism would be harmful without considering how excessive anthropodenial could be harmful as well, including by creating an adversarial dynamic and distorting our understanding of capabilities. - More generally, the framing is misleading. Suleyman concedes in passing that the science is unsettled, yet he presents the view that consciousness requires biology as the general understanding. He also cites academic articles that support this view without citing any academic articles that support taking AI consciousness seriously. Instead, he represents the latter view by citing an AI company, an advocacy organization, and a newspaper op-ed. A reader could be forgiven for forming the impression that experts broadly agree that Microsoft AI is correct and Anthropic is running ahead of the evidence. In fact, Anthropic's consistent acknowledgement of uncertainty is arguably more reflective of the current literature. To be clear, I agree with Suleyman that we need to discuss this issue much more, and I agree on several substantive points as well: that current AI systems are unlikely to be conscious, that over-attribution of consciousness is currently more likely than under-attribution, that over-attribution of consciousness can be harmful, that we should avoid creating plausibly conscious AI systems unless we can do so responsibly, and that we should avoid creating AI systems that appear likelier to be conscious than they are. But there is more to the story. The probability of consciousness in current AI is already nonzero and arguably non-negligible. Under-attribution of humanlike traits to AI systems in general is already a risk, and it can be harmful as well, not only for welfare but also for safety. Even if we should avoid creating plausibly conscious AI systems for now, we should prepare for the possibility that someone will create them anyway. And we should avoid creating AI systems that appear less likely to be conscious than they are, as well. My impression is that Suleyman recognizes these nuances, but thinks that he needs to take a reductive, all-or-nothing stance to persuade the industry to design AI systems in the right kind of way, and to persuade the public to interact with them in the right kind of way. If this is the case, I can appreciate the intention, but I disagree with the strategy. Acknowledging what makes the issue both important and difficult is a precondition for addressing it with clarity, especially since the risks are likely to change rapidly over time.
11
22
101
9,406