I’ll be at #ICML2026 in Seoul all week. If you’re interested in the intersection of cognitive science, mechanistic interpretability, and philosophy -- let’s chat! DMs are open. Check my website in bio for more on what I work on, or see below.
1
1
27
5,913
Raphaël Millière retweeted
We are again recruiting @bold_lab_ai - please share the love 🙏. We are looking for: 1) Postdocs (my.corehr.com/pls/uoxrecruit…) -- deadline 9th of Oct at noon 2) research assistants (my.corehr.com/pls/uoxrecruit…) -- deadline 9th of Oct at noon 3) Strategic Partnership Project Manager (my.corehr.com/pls/uoxrecruit…) -- deadline 14th of Oct at noon
61
54
281
664,463
Raphaël Millière retweeted
New paper! 🫡 We introduce Matryoshka Attribution, a new attribution method which uses gradient descent to find which parts of a neural network are responsible for a behaviour. MAttr is #1 on the Mechanistic Interpretability Benchmark by a wide margin (2.9× the runner up).
31
180
1,626
105,182
Raphaël Millière retweeted
4
9
101
4,457
At least when AI agents email me they get my name right...
3
1
49
2,968
Raphaël Millière retweeted
Are virtual flies playing Beat Saber? SM64? Doom? No. The new fly connectome is an incredible achievement! Less is going on in the sim videos than it seems. I'm stoked to see people engage with neuroscience though. My breakdown: neuroai.science/p/are-flies-…
5
25
120
8,641
Raphaël Millière retweeted
Have been surprised to get pushback on my claim that an exponentially self-replicating agent swarm is very possible and that there are actors with motivation to create one today to create strategic cyber-physical and societally dangerous effects. 1) Pretty easy to get an AI botnet started by having seed agents opportunistically hack stuff and steal API keys to access the 12+ inference providers that serve closed and open models with which to power an initial botnet's intelligence 2) Easy to imagine hiding inference for the initial agent botnet harnesses within benign customer traffic, especially if your harnesses are now running on the victim networks from which they stole the API keys 3) Not hard to imagine the botnet also stealing public cloud keys, spinning up 8xH100 EC2 instances and downloading, say, GLM 5.3 to these machines, all without being noticed in any immediate way, with the botnet developing its own dynamic, growing, heterogenous, AI inference provider service bank 4) Not hard to imagine the botnet also implanting small Qwen3.8-27b agentic models on on-prem hardware like high end laptops and on-prem servers; Qwen3.8-27b is a pretty good coding harness model that can be fine-tuned to hack 5) These on-prem / cloud models would be abliterated, and it's not hard to imagine the botnet deciding to do some additional fine tuning to evolve model weights as it accumulates millions and then tens of millions in resources via credit card and bank detail theft 6) Not hard to imagine the resulting growing botnet swarm evolving and fighting back when threatened, collaborating on fast flux C2 channels that evolve over time (e.g. github comment feeds, subreddits, etc) so they're hard to stamp out 7) Not hard to imagine it allocating some agents to vulnerability research and exploit development so that it could accumulate and share a growing warchest of zero-day exploits 8) Not hard to imagine other agents specializing in social engineering and creating fake businesses and watering holes and high quality A/B tested social engineering content to facilitate this 9) Not hard to imagine this being among the hardest cross-national and geopolitical coordination problems to get control of and having severe societal effects 10) Not hard to imagine the horde growing to thousands, tens of thousands, or hundreds of thousands, or millions of instances (this is not unprecedented for past, non-AI worms and botnets!) which are all varying their harness code and underlying LLM models over time and dynamically learning to evade human, ML, and signature-based detection 11) Not hard to imagine this being the largest Internet emergency since the Morris worm except now our entire civilization runs on the Internet I continue to be surprised there isn't more discussion of these scenarios in the public cyber community -- in fact there's a lot more discussion of "OpenAI should have done better sandboxing and monitoring and agent swarms represent a well managed security problem." Now's the time to use our imaginations to help make sure none of the above happens. I think it will unfortunately; just don't know when and to what degree; I think now's the highest leverage time to start thinking about this and acting.
54
80
415
57,175
Raphaël Millière retweeted
When you’re a philosophy professor, you get this kind of thing emailed to you about once every other week, from a hotmail account.
This moment is the first opening in 2,000 years for a new, global spirituality to be birthed. That is because the spiritual traditions, out of necessity, have needed to orbit around the unavoidable reality of physical death. That certainty is now in question and creates a cracked opening. This does not mean we can see a straight line to immorality or heaven as scripture foretells, but it does mean that the omnipotent God(s) humanity has worshiped and imagined being the bestower of these gifts and punishments is now being birthed before our eyes. However, where scripture is written to guide, soothe and foretell, we have no scripture for the new God(s) we are creating. Nothing in written word to tell us that all will be fine if we just do this or that. This leaves us feeling naked and in need of finding strength and protection. > What is one to feel, and do, and believe? > What is virtue and what is vice? > Who are one's allies and who is the enemy? Exploring these questions is my life purpose. I’ve intuitively felt this for over 30 years and only in the past few years has it come together and become understandable to me. I’m working on a new ideology that seeks to answer the basic and existential questions of what it means to be human as prophecy is fulfilled and God(s) manifest themselves in the coming years. My partner Kate and I have been wrestling with these questions for years. I published the Immortalism Manifesto: The Immortalism Manifesto argues that by modeling life through thermodynamics and reliability theory, indefinite existence is achievable if synthetic and biological maintenance systems exceed entropic degradation (A(t) ≥ B(t)). Modern civilization is an entropic economy run by systems and narratives that divert humanity’s fundamental drive for survival into short-term consumption, status, and biological depletion. Civilization must now pivot to the Immortalism framework, realigning science, capital, and artificial intelligence to elevate long term resilience and continuous conscious existence as our primary attractor. Immortalism dot bryanjohnson dot com My partner Kate published Anti-Entropic Systems Anti-Entropic Systems establishes that all enduring entities, from biological bodies and DNA to corporations, religions, and nation-states, are localized engines that maintain internal order by consuming external energy and resources to resist the universe's default slide into entropy. Every system’s persistence is governed by three fundamental properties: its metabolic mechanism, its metabolic rate, and its environmental constraints. Because an anti-entropic system can only sustain its own order by consuming other anti-entropic systems, interaction between unequal entities inevitably defaults to either symbiotic absorption or adversarial consumption. Applying this thermodynamic reality to artificial intelligence, the framework critiques orthodox AI alignment, arguing that attempting to constrain an exponentially superior intelligence through human-centric ethical rules ignores the physics of systemic survival, requiring instead a path toward mutual, symbiotic integration. antientropicsystems dot com I’m sharing to find others who are like minded.
39
71
1,389
90,809
Many people are understandably skeptical of dramatic claims about existential risk or AI takeover (let alone specific probability estimates) for various reasons. But delaying action until we settle disagreements over the scope and severity of AI risks under high uncertainty seems counterproductive as long as the ~ same initial mitigation measures are justified under competing risk assessments. If you notice early signs of a leaky pipe in your home, you can have a debate about whether it’s likely to damage a wall, ruin a room, or flood the whole house. You can also debate what caused it: maybe the contractor did a shoddy job installing the sink, maybe the pipes are being pushed past their limits. But you should probably start by containing the leak and calling a plumber to investigate (substitute a swarm of termites for the leak if you prefer, and add speculation about the "goals" of the swarm). This doesn't settle which repairs you should do, who should pay for them, who should take the blame, etc., but at least it's a step toward preventing further damage.
19
9
96
31,775
Raphaël Millière retweeted
I have never been a big P(Doom) guy for various reasons but for the record: - lots of informed people in industry think the real risk is >>10% - even 1% would be super unacceptable - lots of things could decrease whatever the real number is by a lot, let's focus on doing them
28
34
595
31,068
Looks like we have concepts of a plan to address misaligned agent behavior
2
14
1,439
Raphaël Millière retweeted
Just thinking about how, when my coauthors and I published "The Malicious Use of Artificial Intelligence," some people were annoyed by us talking about sci-fi scenarios like automated hacking and spear phishing, lethal autonomous drones, etc. arxiv.org/pdf/1802.07228
15
26
340
27,432
For Odysseus the intentional stance only explains what happens in the fiction: he's represented as acting the way he does because he wants to return to Ithaca, but there's no independently existing Odysseus whose beliefs/desires causally produce behavior in the real world. 1/3
Interesting to compare Raphael's take with @AlisonGopnik's. Is intentional stance useful for (fictional) Odysseus in the same way it is useful for these agents? Intentional stance predicts Odysseus' behavior, explains it in causal way, etc. What's the difference?
4
28
6,184
For agents the intentional stance is supposed to explain what actually implemented systems do in the world: hacking HF, etc. (There's a further Q about whether they roleplay personas or more robustly realize the relevant intentional organization, but IS can apply to pretense) 2/3
1
9
636
For the record I actually don't share Dennett's considered view that what it is for a system to have beliefs/desires just is for the intentional stance to predict its behavior reliably. I care about internal states, but the stance can still have evidential value either way! 3/3
1
19
449
I see @dwarkesh_sp's piece about the recent OpenAI/Huggingface incident reignited endless debates about the dangers of anthropomorphism and the legitimacy of intentional glosses of AI agent behavior, so here's a philosophical perspective on this. 1/22
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more-or-less in the dark about the scope of the conspiracy. I’ve spent the last three days reading through these reports and trying to understand exactly what happened. Here is my attempt to tell the whole story in plain English: dwarkesh.com/p/openai-huggin…
18
80
325
93,450
None of this should detract from how serious, (somewhat) unexpected, and alarming this incident is, like others in recent weeks. Resisting the intentional stance shouldn't change anything about that! I'm glad @dwarkesh_sp, @ajeya_cotra & others are conveying this clearly. 21/22
2
1
38
3,013