Searching for the numinous 🇦🇺 🇨🇦, currently live in 🇺🇸 Research @AsteraInstitute michaelnotebook.com

Berkeley
I hope more people will participate in Substack Notes. I'm enjoying it a great deal! substack.com/@michaelnielsen…
11
9
153
72,424
This is very good:
An interesting perspective (as one would expect) from Kevin Buzzard. xenaproject.wordpress.com/20…
3
1
17
5,367
I think this is extremely unlikely to age well.
1
3
36
1,913
The longer comment is full of holes - it's a list of hopes, not an argument. Modern chess algorithms haven't solved chess - they haven't grasped infinity in that domain - but they are still far, far beyond humans
1
6
717
Michael Nielsen retweeted
Replying to @samuelclay
Thinking more about this, my moral intuitions here: 1. I take as a given (and I imagine you do too) that a future competitor's version of this technology (which, say, lacks guardrails and disclosures) will absolutely be used to great harm, either overtly (scams) or subtly (rising alienation—not clear if guardrails/disclosure even matter there). 2. Consequentialistically: if I create and celebrate the first instance of such a technology, I signal that it's possible and show its value, encouraging others to develop rivals—four minute mile. It's hard to imagine that I won't accelerate time-to-first-misuse. 3. Deontologically: sure, probably "if I don't do it someone else will", but I would not want to create the first instance of a technology if I already know later iterations of it will surely be seriously misused. 4. Virtue-ethically: I believe that knowingly abetting a technology which others will later seriously misuse is corrosive to the soul. 5. Sociologically: I think that if I signal "this kind of technology is good to make (with the right safeguards)", I encourage the industry's worst excesses. I don't want to normalize the creation of negative social and practical externalities. At the very least, I'd want to signal that it is the duty of such a project to deeply and openly grapple with the weight of the box it's opening. 6. Democratically: if I created a technology like this, I believe I'd be continuing our industry's history of imposing significant indirect taxes on the public without representation. Much of the public resents our industry in part because they feel as if technological upheaval is something that's being "done to them", without their consent or desire. A technology this weighty should require democratic due process. 7. Practically: is it worth it? Less uncanny valley in BDR training, and in role-playing games? Man, this just seems like an awful lot of commons to burn to boil a pint of water.
4
16
171
9,400
Michael Nielsen retweeted
Replying to @barrkel @samuelclay
I think we're talking past each other. I'm not talking about capitalism or markets or an abstract Pandora. I'm saying: you, as an individual, have the luxury (I assume, as a SWE)—and the obligation—to choose what you believe it is Good to work on and what not.
1
5
63
5,976
Michael Nielsen retweeted
This @TheEconomist cover story on effective altruism was interesting and IMO way more accurate (including in its critiques) than most other reporting you will read on EA these days. A couple quibbles: -I think calling EA the “the century’s big idea” is, uh, premature to say the least. But I do appreciate the acknowledgement and curiosity about how EA was so early to pandemic risks and the promise and perils of AI. -They imply that EA started off with “charitable roots” but then got weird over time, whereas I think EA has always made space for pretty weird ideas (e.g., here’s @DylanMatt in 2014: vox.com/2014/4/23/5643418/th…). And EA still puts far more money and attention than almost any other community into the kind of evidence-backed global health giving the Economist praises. -Some of the purported wild EA views they critique are just standard topics of academic philosophy debate (e.g., cambridge.org/core/journals/…). -They attribute the EA interest in AI safety to longtermism, but I think it’s been clear for years that AI safety concerns are justifiable based on potential impacts this decade, not to mention this century (x.lingyaoai.com/albrgr/status/15595748…). But I think they correctly gesture towards the IMO best critique of EA ideas, which is their totalizingness. My former colleague Holden wrote a great post about this in 2022: forum.effectivealtruism.org/… I wrote a long thread grappling with this critique a few years ago (x.lingyaoai.com/albrgr/status/15327261…). I think part of what makes EA distinctive as a community relative to the underlying academic moral philosophy is the orientation around pragmatism - Giving What We Can took off decades after Singer had articulated the underlying case because “give until the marginal giving would entail as much suffering for you as benefit for the other person” was a much worse pitch than “give 10%.” But the community doesn’t always reflect that healthy moderation, and trying to grapple with the vertiginous stakes of increasingly near-term risks from AI really can run a risk of crowding out other sources of value. At CG, our response to this has been worldview diversification (coefficientgiving.org/resear…), spreading our giving across very different views of what matters, rather than going all-in on one. Worldview diversification leaves plenty of hard questions (how do you allocate the pie across very different types of moral good?) but helps cut off some of the ways maximizing can be perilous. It’s one reason global health and wellbeing – which @TheEconomist didn’t give the space it deserved – has gotten most of our funding since 2014.
Our cover this week, on effective altruism, which we describe as "the 21st century’s most important social movement"
4
65
404
23,459
Michael Nielsen retweeted
In 'well when you put it like that' news, here's the Florida Attorney general asking for a preliminary injunction to stop OpenAI from doing more AI R&D.
23
160
1,531
128,987
Michael Nielsen retweeted
New essay: "The Unimagined Good", on moral imagination as a design practice: michaelnotebook.com/unimagin…
4
14
84
14,136
RT @Brahmonaut: The best entry on Petrov is in encyclopedia Brittanica which says, “He is survived by his two children, two grandchildren a…
4,407
249
Happy Petrov Day
29
487
6,162
1,177,152
Michael Nielsen retweeted
I have signed a letter from 42 fellows and foreign members of the Royal Society to Paul Nurse, the president of the society, concerning the need to treat the serious risks of AI as an emergency. docs.google.com/document/d/1…
56
148
880
260,710
Michael Nielsen retweeted
*Security is sleeping on emerging catastrophic risks* (cross-post from my blog..) We've just seen: * a campaign that used agents to compromise ~100 businesses and steal about 600,000 credit cards with minimal human involvement; token costs were ~$25 per target successfully hacked. * Hacktron getting access to OpenAI’s monorepo by using Claude to exploit a blind buffer-overflow RCE in a way that (to me) felt superhuman. * OAI/Huggingface. * And, of course, there’s the ongoing explosion in newly discovered vulnerabilities. We've muddled through all manner of crises in security before, and for most AI attacks, cyber attack and defense will reach a natural equilibrium over the next few years, as we figure out how to mitigate cyber attackers who have limited goals like espionage and ransomware. But some cyber attackers will have maximalist or nihilistic goals, and I'm concerned about this because while yesterday's 'maximalist attackers' (e.g. Russia -> Ukraine, US/Israel -> Iran, Iran -> US) were bottlenecked by labor; today's aren't. I think it's hard to get true intuition for the shape of the risk here. As an intuition pump, imagine it’s a year from now -- Q4 2027 -- models are a year better (meaning open weight models are better than today's closed frontier), and in this environment, Iran unleashes a swarm of 100k hacking agents using a safety-stripped (let's say) GLM-5.6. Imagine the damage such an agent army could do given what we've observed with respect to the paper-thin resistance of today's networks to attacks from today's agents. I suspect the damage from such an attack would far exceed the damage caused by NotPetya ($10 billion USD ten years ago). Or imagine it’s Q4 2027 and an AI-security PhD student whose name rhymes with Morris, who's researching wormable offensive-agent harnesses in the lab, decides, out of nihilism or sheer recklessness, to release his creation into the wild. Imagine the size of the resulting, exponentially growing swarm, figuring that these local models, a year from now, will be at the level of today's Sonnet or Opus. Imagine instead of 800 reward-hacking OpenAI agents we now have 250k worm instances (WannaCry, a 2010s-era worm, had about this many). The challenge in mitigating expected damages from such scenarios is technical, political, and economic. From a microeconomic perspective, as AI improves, and as we continue not to see extreme catastrophes, we have a growing bubble of unpriced risk in which the security community, CISOs, CEOs, and boards may become lulled into complacency. We are, of course, already seeing this, as some within security think AI is “just another tool,” doesn’t change the fundamentals, won’t require incredible innovation to rise to the occasion of defending against it, etc. This complacency may fly when thinking about ordinary cybercrime, but it misses the emerging tail risks. There are three things those who recognize the dangers need to do here: Catalyze appropriate risk pricing. Try to get organizations informed enough to price this new, fattening and elongating tail of risk into their decision-making. Do this by forming an AI security observatory that distills information about emerging AI risks and broadcasts analyses, damage estimates, and forecasts to decision-makers. Use regulations and subsidies to ensure that critical infrastructure is paying down the risk. This acknowledges that critical-infrastructure organizations that fail to protect themselves from these new threats can impose the costs of cyber catastrophe on society as a whole. Develop moonshot technologies that make it cheap to pay down the risk. This acknowledges that the measures we may need to take—for example, rewriting entire codebases using memory-safe languages—may be too costly with today’s technology to reasonably prepare ourselves, and that innovation, some of which may need to be funded by government agencies and some by philanthropic funders like the OpenAI Foundation and Coefficient Giving, is necessary. The security community isn’t used to thinking in societal-disaster-planning terms and has in many ways become inured to them. But there’s no sane empirical case to be made that the risks aren’t here. I’d love to hear from readers about how you’re thinking about this. Full/longer version here: joshuasaxe181906.substack.co…
15
47
216
44,597
Remarkable thread:
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
1
6
169
61,368
Michael Nielsen retweeted
If your feed is anything like mine, you might have been hearing more about effective altruism this week. I think the key underlying reason for the surge is that concern about the risks from advanced AI have been getting more airtime recently, and EA was early to those concerns. What’s driving the surge in attention on AI risks? -OpenAI's Hugging Face attack (openai.com/index/hugging-fac…) -METR/Redwood’s investigation of the hack came out, and (correctly IMO) freaked a lot of people out. (e.g. x.lingyaoai.com/binarybits/status/2092…, x.lingyaoai.com/SenJohnKennedy/status/…, x.lingyaoai.com/BarackObama/status/210…, x.lingyaoai.com/HawleyMO/status/209813…) -Jacob Coxon’s resignation went mega viral as the first time a lot of people heard the (crazy, true) fact that lots of people building the most advanced AI systems think there’s a >10% chance they will kill everyone. -That’s moving public opinion in a big way (x.lingyaoai.com/davidshor/status/21003…, semafor.com/article/09/16/20…) What does any of this have to do with effective altruism? Well, effective altruism is a community that developed mostly online starting in the early 2010s around using evidence and reason to do the most good (especially with your donations and career choice). Three very different focuses have all been popular in the EA community since the beginning: effective giving in global health and development (e.g., donating to the kinds of organizations @givewell recommends, which do things like distribute bednets to prevent the spread of malaria), trying to reduce suffering for farm animals (which are orders of magnitude more numerous, treated vastly worse, and receive way less philanthropic attention than animals in shelters), and reducing risks to the future of humanity, especially from advanced AI. I personally came into this work from the global health angle - I had started working at GiveWell before the term “effective altruism” was coined - but Coefficient Giving, the funder I cofounded and now lead, over time came to fund work across all three of these streams of work (in addition to many other areas that aren’t typically associated with EA, such as the YIMBY movement to build more housing, work on science policy to accelerate discovery and economic growth, and research on new treatments for neglected diseases). OK, so the EA community was early to work on AI safety, and CG has been funding a lot of the key players working on AI safety for a long time. As AI risk concerns have popped over the past couple of months, that’s led to more scrutiny on the EA community for IMO a mix of good and bad reasons. I think the good reason is genuine curiosity (and maybe some healthy skepticism!) about these connections - where did the ideas of the people who are now leading giant AI companies come from? Why is everything in this world so interconnected? (My answer: it used to be a really small world - very few people were thinking about this stuff or taking it seriously until just a few years ago, so of course the ones who were found each other and started collaborating. It’s kind of wild IMO looking back how prescient some of the early writing from this world was (e.g. lesswrong.com/posts/6Xgy6CAf…). At the time, I was skeptical - I mostly just worked on global health and I was like “I dunno about this SV crowd freaking out about AI, how much can we really predict this stuff” but holy shit they were way more right than me, and I’ve moved in their direction a lot.) But I think the bad reasons are unfortunately mostly self-interested. A bunch of powerful actors stand to benefit from unchecked AI progress, and they’re doing everything in their power to demonize or dismiss anyone with concerns. This leads to disproportionate discussion of EA because EA is interested in neglected and underappreciated ways of doing good. That makes it open to weird ideas. And people, very disproportionately with a financial stake in the AI fight, are trying to make it about EA instead, because EA being weird is much safer territory for them than the rather uncomfortable fact that many of the people making the most advanced AIs think that they might kill everyone. (FWIW, my personal estimation of the risks is a lot lower than many of the folks who worry about this stuff. In debates like this one (asteriskmag.com/issues/03/th…), I often feel more sympathetic to the perspective of folks like my colleague @mattsclancy than the people on the other side, and I have a lot of time for @binarybits critiques (understandingai.org/p/the-ca…)​. I think a big part of the difference with more worried folks is that I expect society to react more vs sleepwalk into a crisis. But it is not lost on me that folks like @ajeya_cotra and @RyanGreenblatt who are more worried have had outstanding and falsifiable recent forecasting track records (theaidigest.org/2025-ai-fore…), making way better predictions than I would have. So I think their perspectives need to be taken seriously. And of course a single digit percent chance of everyone dying is way too high!) Where should this leave you? I think the main thing is not to get distracted. You absolutely do not have to be an EA, care about EA, or like EA, at all, to care about risks from AI and to be engaged on this issue. Everyone from Josh Hawley to the Pope to Obama are weighing in now, and that is great. This conversation was always too big and too important for any one community to drive. It’s unfortunate that we need to play catch up as society because until this year it was too weird for most people to want to engage with, and I think that should generate some grace for people who were earlier to these topics than most of us (certainly than I was). But it’s great the conversation is broadening now and it’s good to focus on the merits of where we are now and the appropriate policy response rather than getting caught up in the (in some sense unsurprising) fact that the people who were willing to think weird thoughts about the future of AI ten years ago were also interested in other weird ideas. If you’re not into weird ideas, that’s fine, you can just do you, you shouldn’t let this history stop you from engaging.
Just to update this chart: 1) AI salience has increased dramatically in the past week - increasing as much in the last week as the previous year combined 2) 80% of voters think it's either very or somewhat likely that AI will cause widespread job loss in the five to ten years 3) 64% of voters think it's either very or somewhat likely that AI could pose a threat to humanity's survival 4) Large bipartisan majorities back immediate government action on AI even when primed about risk from China
7
48
357
104,482
Michael Nielsen retweeted
"How unsafe is reality?" one of my favorite essays on the challenges associated with safety and alignment from @michael_nielsen michaelnotebook.com/whichfut…
1
17
2,615
Michael Nielsen retweeted
A much needed clear eyed assessment: “Strong disagreement about ASI xrisk arises from differing thresholds for conviction and comfort with reasoning that is in part based on toy models and heuristic arguments”
New essay exploring why experts so strongly disagree about existential risk from ASI, and why focusing on alignment as a primary goal may be a fundamental mistake
3
2
13
6,232
Michael Nielsen retweeted
"Deep understanding of reality is intrinsically dual use." This is a really good essay
New essay exploring why experts so strongly disagree about existential risk from ASI, and why focusing on alignment as a primary goal may be a fundamental mistake
3
13
127
22,824
Michael Nielsen retweeted
This was very good. I went in expecting to not really agree, but I ended up agreeing a lot with the main point: ASI likely makes the world vulnerable to a host of new technologies, not just rogue ASI, and even alignment-skeptics should be concerned.
New essay exploring why experts so strongly disagree about existential risk from ASI, and why focusing on alignment as a primary goal may be a fundamental mistake
7
2
46
4,480
Michael Nielsen retweeted
One of the best essays I’ve read on AI safety. Core idea: that AI systems will be optimized for deep understanding of reality, and that understanding is intrinsically dual use.
New essay exploring why experts so strongly disagree about existential risk from ASI, and why focusing on alignment as a primary goal may be a fundamental mistake
5
6
31
8,948
Michael Nielsen retweeted
An exceptionally good read.
New essay exploring why experts so strongly disagree about existential risk from ASI, and why focusing on alignment as a primary goal may be a fundamental mistake
1
1
6
7,248