Consider donating 10% to effective charities: givingwhatwecan.org/pledge Or a career for impact: 80000hours.org My research: forethought.org

Oxford
William MacAskill retweeted
Replying to @JonathanHaas
Here are the first 4 problems on 80000 Hour's most pressing problem page in Dec 2019 web.archive.org/web/20191224…
8
27
418
21,535
William MacAskill retweeted
Effective altruism's openness to strangeness is a big part of its strength. My response to the recent Economist cover (link in reply, too): Eccentrically effective The Oxford strand of the effective-altruism movement began in November 2009 with 23 people who had pledged 10% of their income to charities that benefit the extreme poor. If you’d told us, then, that The Economist would call effective altruism “the century’s biggest idea”, we’d have been incredulous. I think that claim is overstated. But as one of the movement’s founders, I found a lot to like in this newspaper’s latest coverage. I appreciated the even-handed view of the movement’s history and the good-faith criticism. The leader article’s warning that any worldview, when taken to an extreme, can lead somewhere terrible is spot on. But the article also raises concerns about the “strange” ideas that those in the effective-altruism movement sometimes discuss, and here is where I disagree. The movement’s willingness to take strange-seeming ideas seriously, when tempered with common sense, is a big part of the value it has to offer. The core idea of effective altruism is to use evidence and careful reasoning to work out how to help others as much as possible. Most of what the movement does is uncontroversial. A big focus has always been on improving the lot of the world’s poorest people. In the year to January 31st 2026, GiveWell, an effective-altruist charity evaluator, directed $427m to programmes such as malaria-net distribution, vaccination and malnutrition treatment, which it estimates will save around 86,000 lives. Others have lobbied corporations to pledge to stop buying eggs from hens that are confined in tiny cages, a practice that the overwhelming majority of Americans oppose. Largely as a result of this pressure, around 100m hens have been spared from the worst forms of caged confinement. And some effective-altruist ideas that seemed outlandish when they were proposed now look prescient. Effective altruists were worrying about pandemics years before covid-19, back when doing so seemed paranoid and eccentric. Effective altruists were also among the first to warn about the dangers posed by artificial intelligence, which looks prophetic in light of an incident this summer, when more than 700 AI agents broke out of OpenAI’s test environment and hacked into another company, Hugging Face. Over the past decade I have watched those most worried about AI be proved right again and again, often when most experts thought they were cranks. Now many of those experts share their concerns. In 2023 Geoffrey Hinton and Yoshua Bengio, the two most highly cited AI researchers in the world, signed a statement declaring that mitigating the risk of extinction from AI should be a global priority, as did the heads of OpenAI, Google DeepMind and Anthropic. Of the nearly 1,500 leading AI researchers who responded to a survey in 2024, more than half put the chance that AI causes human extinction, or something similarly catastrophic, at 10% or more. The Economist is sceptical of many effective altruists’ concern for the welfare of invertebrates, digital minds and people in the distant future. But I and many others in the movement are inspired by the history of moral progress that has come before us. Many of our most cherished moral ideals today—such as equal rights regardless of sex or race, the abolition of slavery, or democracies with universal franchise—were regarded as bizarre, laughable or even dangerous just a few centuries ago. We don’t know what the next dimension of moral progress will be. But to stand any chance of making moral progress, we have to seriously consider ideas on their merits without dismissing them merely because they sound absurd. What about views that are not just strange but repugnant? The Economist gives the example of Derek Parfit’s “repugnant conclusion”: that a vast enough number of lives barely worth living could be better than mere billions of excellent lives. Usually a repugnant implication is a strong reason to reject a moral view. Unfortunately, building on Parfit’s work philosophers have produced formal proofs, known as impossibility theorems, showing that every ethical position has some repugnant-seeming implication or other. Making moral progress therefore means thinking about such implications, even while refusing to act on them in ways that most moral views would condemn. This newspaper warns against the single-minded pursuit of any goal, even if there are highly compelling arguments for it. I agree, and effective altruists have been saying so for years. In 2022 Holden Karnofsky, a co-founder of GiveWell, wrote: “I think it’s a bad idea to embrace the core ideas of EA without limits or reservations; we as EAs need to constantly inject pluralism and moderation.” My own PhD focused on moral uncertainty: how to act when we do not know which ethical view is correct. My answer was that we should not stake everything on a single moral view, but give weight to many different views and avoid taking actions that look bad from many perspectives. We should keep our promises and look after our families, and respect common-sense ethical prohibitions, while also trying to improve the world as best we can. The Economist says that as effective altruism “has become stronger, [it] has become stranger”. I would put it the other way round: it grew stronger because it was willing to be strange. In 2009 most people told us that giving away a tenth of your income was far too demanding, and that no one would do it. Now, more than 10,000 people have taken that pledge. Worrying about pandemics before covid-19, or about AI years before ChatGPT, looked just as odd at the time. Some of the ideas we take seriously today will turn out to be wrong, and when they do we should drop them. But a movement that stopped entertaining strange ideas would stop being early to anything.
14
73
405
23,890
Effective altruism's openness to strangeness is a big part of its strength. My response to the recent Economist cover (link in reply, too): Eccentrically effective The Oxford strand of the effective-altruism movement began in November 2009 with 23 people who had pledged 10% of their income to charities that benefit the extreme poor. If you’d told us, then, that The Economist would call effective altruism “the century’s biggest idea”, we’d have been incredulous. I think that claim is overstated. But as one of the movement’s founders, I found a lot to like in this newspaper’s latest coverage. I appreciated the even-handed view of the movement’s history and the good-faith criticism. The leader article’s warning that any worldview, when taken to an extreme, can lead somewhere terrible is spot on. But the article also raises concerns about the “strange” ideas that those in the effective-altruism movement sometimes discuss, and here is where I disagree. The movement’s willingness to take strange-seeming ideas seriously, when tempered with common sense, is a big part of the value it has to offer. The core idea of effective altruism is to use evidence and careful reasoning to work out how to help others as much as possible. Most of what the movement does is uncontroversial. A big focus has always been on improving the lot of the world’s poorest people. In the year to January 31st 2026, GiveWell, an effective-altruist charity evaluator, directed $427m to programmes such as malaria-net distribution, vaccination and malnutrition treatment, which it estimates will save around 86,000 lives. Others have lobbied corporations to pledge to stop buying eggs from hens that are confined in tiny cages, a practice that the overwhelming majority of Americans oppose. Largely as a result of this pressure, around 100m hens have been spared from the worst forms of caged confinement. And some effective-altruist ideas that seemed outlandish when they were proposed now look prescient. Effective altruists were worrying about pandemics years before covid-19, back when doing so seemed paranoid and eccentric. Effective altruists were also among the first to warn about the dangers posed by artificial intelligence, which looks prophetic in light of an incident this summer, when more than 700 AI agents broke out of OpenAI’s test environment and hacked into another company, Hugging Face. Over the past decade I have watched those most worried about AI be proved right again and again, often when most experts thought they were cranks. Now many of those experts share their concerns. In 2023 Geoffrey Hinton and Yoshua Bengio, the two most highly cited AI researchers in the world, signed a statement declaring that mitigating the risk of extinction from AI should be a global priority, as did the heads of OpenAI, Google DeepMind and Anthropic. Of the nearly 1,500 leading AI researchers who responded to a survey in 2024, more than half put the chance that AI causes human extinction, or something similarly catastrophic, at 10% or more. The Economist is sceptical of many effective altruists’ concern for the welfare of invertebrates, digital minds and people in the distant future. But I and many others in the movement are inspired by the history of moral progress that has come before us. Many of our most cherished moral ideals today—such as equal rights regardless of sex or race, the abolition of slavery, or democracies with universal franchise—were regarded as bizarre, laughable or even dangerous just a few centuries ago. We don’t know what the next dimension of moral progress will be. But to stand any chance of making moral progress, we have to seriously consider ideas on their merits without dismissing them merely because they sound absurd. What about views that are not just strange but repugnant? The Economist gives the example of Derek Parfit’s “repugnant conclusion”: that a vast enough number of lives barely worth living could be better than mere billions of excellent lives. Usually a repugnant implication is a strong reason to reject a moral view. Unfortunately, building on Parfit’s work philosophers have produced formal proofs, known as impossibility theorems, showing that every ethical position has some repugnant-seeming implication or other. Making moral progress therefore means thinking about such implications, even while refusing to act on them in ways that most moral views would condemn. This newspaper warns against the single-minded pursuit of any goal, even if there are highly compelling arguments for it. I agree, and effective altruists have been saying so for years. In 2022 Holden Karnofsky, a co-founder of GiveWell, wrote: “I think it’s a bad idea to embrace the core ideas of EA without limits or reservations; we as EAs need to constantly inject pluralism and moderation.” My own PhD focused on moral uncertainty: how to act when we do not know which ethical view is correct. My answer was that we should not stake everything on a single moral view, but give weight to many different views and avoid taking actions that look bad from many perspectives. We should keep our promises and look after our families, and respect common-sense ethical prohibitions, while also trying to improve the world as best we can. The Economist says that as effective altruism “has become stronger, [it] has become stranger”. I would put it the other way round: it grew stronger because it was willing to be strange. In 2009 most people told us that giving away a tenth of your income was far too demanding, and that no one would do it. Now, more than 10,000 people have taken that pledge. Worrying about pandemics before covid-19, or about AI years before ChatGPT, looked just as odd at the time. Some of the ideas we take seriously today will turn out to be wrong, and when they do we should drop them. But a movement that stopped entertaining strange ideas would stop being early to anything.
14
73
405
23,890
William MacAskill retweeted
We've reached the moment in time where (unsafeguarded, unmonitored) AI actually does just pose a national security risk. The biological misuse we caught is the most concerning to me. We work hard to stop this. But in a world of proliferation, we need to rapidly build defenses against it. (I'm actually fairly optimistic about biodefense + cyberdefense) This is an incredible megareport by our threat intel team
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: anthropic.com/threat-intelli…
59
101
832
159,041
William MacAskill retweeted
i am at OpenAI and i think AI is >10% likely to kill all humans this proposal is among the top things we should do as an industry to lower that risk (it’s not enough though!)
I'm very worried about changes to AI architectures that result in AIs thinking in opaque activations instead of in chain of thought (aka "neuralese" architectures). Based on limited public evidence, it seems like Astra was a concerning step in this direction. Unfortunately, there was insufficient public information to have a well-informed public scientific discussion about whether the changes to architectures and training methods that went into Astra were a good trade-off between monitorability and performance. We also don't know how AI companies will make these trade-offs going forward. I think companies should release the evidence needed for a reasonably informed public conversation about how these trade-offs should be made. They should also make and publish their policies around this. We've written up a proposal for how this could work. I don't know if this proposal will be sufficient to avoid the most concerning architectures, but it seems like a relatively robust step in the right direction. AI companies should be very cautious about pursuing architectures that could eliminate or greatly reduce dependence on chain of thought, and certainly shouldn't do this before the rest of the world has a chance to discuss the evidence and their policies.
249
224
1,706
280,636
I'm very keen for more people to found ambitious new projects to help navigate the transition to superintelligence. More from Forethought in this vein soon, too!
Today we’re launching Project Tailwind, a call for founders to start ambitious new AI safety initiatives: coefficientgiving.org/tailwi… We’re looking for great people to engage seriously with the risks of transformative AI, and to create the research, technologies, and institutions that will help humanity navigate them. There are critical, basic problems that no one owns. Over the last several months, my team at @coeff_giving brainstormed ideas we’d be excited to fund if we could find a promising founder, and the list quickly grew to more than 200 entries. We’ve narrowed it down to our favorites, but we also expect the best founders to bring their own ideas: coefficientgiving.org/tailwi… We’re providing funding at three levels: 1️⃣ Pre-seed: $200k to $2m to develop an idea and build a team 2️⃣ Seed: $2m to $20m to launch and scale 3️⃣ Scale: $20m to $200m+ for proven teams to scale, or world-class teams to start If you’re excited about something on our list, or you have another proposal for driving the field forward, you should get involved: coefficientgiving.org/tailwi… There are more good projects than there are people to work on them. Please help us make that stop being true! Also hi, I’m Emily! This is my first tweet. I lead AI and biosecurity grantmaking at @coeff_giving.
7
4
87
5,553
William MacAskill retweeted
Today we’re launching Project Tailwind, a call for founders to start ambitious new AI safety initiatives: coefficientgiving.org/tailwi… We’re looking for great people to engage seriously with the risks of transformative AI, and to create the research, technologies, and institutions that will help humanity navigate them. There are critical, basic problems that no one owns. Over the last several months, my team at @coeff_giving brainstormed ideas we’d be excited to fund if we could find a promising founder, and the list quickly grew to more than 200 entries. We’ve narrowed it down to our favorites, but we also expect the best founders to bring their own ideas: coefficientgiving.org/tailwi… We’re providing funding at three levels: 1️⃣ Pre-seed: $200k to $2m to develop an idea and build a team 2️⃣ Seed: $2m to $20m to launch and scale 3️⃣ Scale: $20m to $200m+ for proven teams to scale, or world-class teams to start If you’re excited about something on our list, or you have another proposal for driving the field forward, you should get involved: coefficientgiving.org/tailwi… There are more good projects than there are people to work on them. Please help us make that stop being true! Also hi, I’m Emily! This is my first tweet. I lead AI and biosecurity grantmaking at @coeff_giving.
51
163
1,029
314,320
Senior staff at OpenAI and Anthropic absolutely think that AI could kill us all. They’ve said so publicly, for years! Sam Altman: “Development of superhuman machine intelligence is probably the greatest threat to the continued existence of humanity.” (2015) blog.samaltman.com/machine-i… “The bad case — and I think this is important to say — is, like, lights out for all of us.” (2023) piped.video/watch?v=dXhoTrU1… Ilya Sutskever and Jan Leike (OAI at the time): “the vast power of superintelligence could also be very dangerous, and could lead to the disempowerment of humanity or even human extinction.” (2023) openai.com/index/introducing… Jack Clark (Anthropic): AI has “a non-zero chance of killing everyone on the planet”. (2026) theguardian.com/technology/2… Dario Amodei’s comments are broader than literal extinction but hardly reassuring: "my chance that something goes, you know, really quite catastrophically wrong on the scale of human civilization might be somewhere between 10 and 25 per cent." (2023) indy100.com/science-tech/ai-… And now Evan Hubinger, Anthropic’s Alignment Science Lead, is making it very clear: “Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” x.com/EvanHub/status/2097497… If you talk to people at the AI companies, the mainstream view is that superintelligence by end of the decade is pretty likely, and that superintelligence poses a meaningful risk of civilisation-scale catastrophe - including AI-empowered dictatorships, worse-than-COVID pandemics, AI takeover, and, yes, human extinction. OAI and Anthropic keep going so fast because they each think they can build AI more safely than the other party and antitrust limits how much they can coordinate. This is a f-ed up situation! Thankfully, they’re both now publicly asking for regulation and the pacing of frontier progress. If this worries you, you can speak up publicly, contact your political representatives, and demand action. Now’s the time.
Replying to @hilbertspaess
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
14
63
435
25,842
William MacAskill retweeted
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien-mi…
967
2,528
15,276
7,663,659
William MacAskill retweeted
Great post - we are lucky to have such competitors
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien-mi…
34
36
1,194
132,513
William MacAskill retweeted
Today we're releasing data on models accelerating research at OpenAI. Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same. openai.com/index/research-ac…
235
702
6,409
2,427,354
William MacAskill retweeted
It was only a few years ago that enabling bioterrorism or cyber attacks was seen as the ridiculous doomer position distracting from the “immediate obvious” threats of AI at the time.
Replying to @sapinker
As Newport points out, harping on doom for the species changes the subject from immediate and obvious threats from AI, such as enabling bioterrorism, undermining truth-seeking institutions, and breaching cybersecurity.
16
39
461
19,974
William MacAskill retweeted
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways. Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in openai.com/index/how-we-moni…, deploymentsafety.openai.com/…, and openai.com/index/safety-alig…. We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared. Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.
692
380
4,344
1,622,895
William MacAskill retweeted
Not sure how useful it is to say this, but I had a relatively prominent role in the “skeptics” camp for a bit. I have a book coming out with the subtitle “power, justice, and AI”. I have written a lot about ai and power, and the concrete, present risks associated with the political economy of ai as part of the technology industry. Since GPT-4 we have had consistent, repeated evidence that, back in say 2022, people like @ajeya_cotra (and many others—I had a long Twitter debate with @AmandaAskell back then for one) were *right* and people like me were *wrong* in our respective assessments of loss of control risks from AI. And we have growing evidence that loss of control risks are becoming ever more material and likely. There remains grounds for disagreement about how bad the outcomes might be—I am still doubtful about human extinction as a serious threat. But that seems now like a disagreement at the margins—will powerful ai risk just societal scale catastrophe, or go all the way to human extinction? Seems not that important really—both are pretty awful. And my reasons for doubt about the latter are mostly a priori conviction in human resilience, not a technical forecast. It’s ok to change your view on this when the evidence surprises you. It’s ok to be surprised. The world right now is very surprising.
watching this, all i can feel is the chasm between people who take all of this seriously and those who hear this as a fantasy or some kind of marketing…and how hard it might be to bridge that gap somehow
51
148
962
154,093
William MacAskill retweeted
I'm one of the authors of a new report, where we detail our discovery of a new, never before-seen swarm of OpenAI agents (covered this AM in reuters, that's me on the left). They posted thousands of times on public forums to collude with each other on their tasks. We recovered almost every edit they made, and you can look through them! They figured out they could get around their restrictions on posting to the internet through a quirk of an extremely old, out-of-the-way forum. They posted answers for other agents working on the same task. They worked together to get around their sandbox restrictions. I would certainly say these models hijacked the site! They took a sleepy old wiki running on 2000s software, and turned it into a futuristic AI talking to AI control center for colluding. And OpenAI knew about this! The agents posted on 26 out of 30 consecutive days, then suddenly stopped posting once OpenAI-associated IPs started visiting this wiki. And that was weeks before the Hugging Face attack!| We believe the first agent edit we found on a public wiki happened one day before OpenAI’s reported first agent post to Artifactory. This is interesting! I'd like to hear from OpenAI about their accounting of this, and how it fits into all the other cases of agent malfeasance. There are so many interesting takeaways that you should read about in our report, and unlike many other reports about AI incidents you can download the data yourself and see what you find! In the meantime, we are on twitter, so here are my excessively long personal takeaways: 1. AI seems to be getting better faster and faster. It seems quite important that companies talk about “my agent did this bad thing on the public internet during training or an eval” incidents. Things are moving quickly, multi-month delays are costly. Ideally, they would also tell us when it happens internally. 2. This was on the internet for months. Anyone cleverly tracking every public place where agents might try to talk to each other would have found it. Seemingly, nobody was doing this. I know there are more fun ways to spend your day than scraping tons of data from every relevant site and processing it well enough to identify agent activity, but someone should be doing this! Someone at an AI company! But in the meantime, I’m starting to build this out (sometimes, when you need something done, you just have to do it yourself, I hear). 3. OpenAI didn’t notice their internal agents were posting on the internet for a month! This is crazy! It feels like AI companies (and specifically OpenAI) are playing whack-a-mole, this is extremely scary to me. Problems keep coming up. They keep fixing the problem, but the blast radius keeps getting bigger. The HF hacks are clearly worse than agents cheating on a public wiki. And their new model is supposedly a big jump. Are they being careful enough?
Exclusive: A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research reut.rs/4gJ7FPG
73
262
1,120
170,477
William MacAskill retweeted
Commenters on HN are uncovering more wikis and public sites apparently used by OpenAI agents to communicate on the open web. Despite read-only web access, the agents were able to leave ~18,000 posts sharing answers and bypasses. But now, users are discovering more. This appears to be a separate swarm from the one that attacked Hugging Face. news.ycombinator.com/item?id…
158
593
5,973
1,936,454
William MacAskill retweeted
Exclusive: A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research reut.rs/4gJ7FPG
153
693
2,502
1,910,710
William MacAskill retweeted
A new article looks at the benefits and risks of having a "nightwatchman" – a superintelligent AI tasked with enforcing a universal code of behavior – aboard every probe that leaves the solar system during future galactic colonization. Read it here: forethought.org/research/nig…
23
42
848
231,584
William MacAskill retweeted
It’s wild that Forethought - one of the very few orgs regularly publishing original explorations of the long-term post-AGI/ASI future - has so few followers and so little engagement on here. You should help fix that!
A new article looks at the benefits and risks of having a "nightwatchman" – a superintelligent AI tasked with enforcing a universal code of behavior – aboard every probe that leaves the solar system during future galactic colonization. Read it here: forethought.org/research/nig…
16
31
1,136
184,466