Reporter covering AI @WSJ | Formerly @WIRED, @TechCrunch, @markets | DM me off the record on Signal @ mzeff.88

San Francisco, CA
Pinned Tweet
Personal news! I'm joining @WSJ to cover AI. I've had an amazing time at WIRED working with brilliant colleagues across the newsroom, but I'm very excited for this next chapter. I'll continue to cover OpenAI, Anthropic, and the bustling AI industry. I start next week!
122
27
1,125
52,292
Quite the essay. “Perhaps I should have stayed and fought for fundamental shifts in our staffing and culture, but in practice, my colleagues and I were so busy sprinting that we seldom had the chance to consider big changes, much less to actually make them.”
.@dgrobinson, who worked on OpenAI’s safety team, resigned from the company this week. “I believe we need to look deeper than specific rules or new laws. We need to talk about culture,” he writes: theatlantic.com/technology/2…
2
1
22
1,740
A 20-something Harvard dropout built Facebook into a social media giant. A 20-something MIT dropout is helping usher it into the AI era. My profile of Alexandr Wang, the camo-wearing Gen Z billionaire behind Meta’s new consumer agent app, featuring ~those~ memes wsj.com/tech/ai/alexandr-wan…
10
17
168
26,423
Max Zeff retweeted
Update: Our story now names the three people OpenAI fired for alleged misconduct, including the sharing of company info with a third-party AI-safety org. Jasmine Wang, Tomek Korbak, and Mikita Balesni.
2
14
67
6,121
NEW: OpenAI has parted ways with three researchers that allegedly shared sensitive information with a third party AI safety organization. Developing... wsj.com/tech/ai/openai-parts…
27
58
435
121,715
Update: Our story now names the three people OpenAI fired for alleged misconduct, including the sharing of company info with a third-party AI-safety org. Jasmine Wang, Tomek Korbak, and Mikita Balesni.
2
14
67
6,121
Great scoop also, it’s remarkable how much the national discourse around AI regulation has changed since greg initially made this commitment in August 2025. Don’t think many expected it would become quite this big an issue so quickly
NEWS. More AI fallout ahead of the midterms. Greg Brockman, the OpenAI co-founder, is no longer making the second $25 million donation promised to the super PAC Leading the Future. Scoop with @MikeIsaac. nytimes.com/2026/09/30/techn…
1
21
4,734
big day for minecraft
Tonight. 8 p.m. My Minecraft City. Subscribe now: piped.video/@miket757
1
3
1,025
So much happened in this press conference
Trump: This guy (Sundar Pichai) is a monster, and nobody knows him. What a great life! To be a Monster and not have to go through this. All you have to know is — sundarr! Sundar is his name.
1
11
3,477
Max Zeff retweeted
It’s really incredible how there is now a type of guy in the discourse whose sincerely held take is “we need to stop worrying about these abstract sci-fi risks and focus on present-day harms like emergent ecologies of digital minds engaging in unauthorized hacking.” Bravo.
46
82
1,368
90,140
👀
Hi, Tomorrow we are re-opening the Pro $200 subscriptions to new subscribers, but together with it we are also changing how we calculate the usage for it. In effect, if you do the math, it will net out at half the dollar in API spend compared to the old Pro $200 plan. Now that it's said, let me explain why this is happening and why you will still get more work done than if you were on the Pro $200 subscription one month ago. (a) We didn't want to compromise in other ways and are committing to not reintroducing the 5h limit, so that you can fully use the weekly usage when you want. (b) On the subscription, we guarantee that over time you always get more work done and with an increasing level of quality. This means that you will continue to get more value per dollar spent as a result of models getting more efficient and us passing down the improvements in the form of API price reductions. (c) We don't want to put an incentive on ourselves to artificially inflate the API list prices to make it look like you are getting a lot (and workaround it through discounts, etc). Instead we want to continue to both rapidly reduce prices and increase capabilities of models on the API. This week we introduced GPT-6 Sol and GPT-6 Luna at 50% of their previous price. Over time, we see prices go low enough that it makes sense for most to buy usage as needed without there being a significant gap between what you get in a subscription and what you get in the API for a dollar spent. (d) Tomorrow, we are adding more things to the subscription that won't draw on the usage, I won't reveal what that is yet. I wanted to be transparent before all the big announcements tomorrow. Lots of new exciting things are coming to the subscriptions that will make it super compelling, but I wanted to make sure to share this change ahead of time so you can all understand it before we shower you with good news. Codexingly, Tibo
3
1,549
New: OpenAI says it's scrapping the release of its new AI model, GPT-6.1 Astra, over safety concerns. The company originally expected to release the model in October, but now says it's pushing ahead with future model generations, and ensuring they're safe wsj.com/tech/ai/openai-chatg…
7
28
125
83,032
Compared to previous OpenAI models, GPT-6.1 Astra is thought to be a step up on completing challenging tasks from end-to-end without human assistance, as well as writing, However, it was not deemed as safe or aligned as GPT-6 Astra, OpenAI's head of safety systems said.
3
5
1,460
The decision comes at an interesting moment, after a summer of concerning AI security incidents and right as public officials are paying more attention than ever to AI safety. It's also just ahead of devday tomorrow, where openai is shipping a ton of stuff to win over devs.
5
1,127
Max Zeff retweeted
Over the last few weeks, so many people have asked me: if people at Anthropic believe that AI has a meaningful chance of killing us all, why are they building it? @berber_jin1 and I did a deep dive into Anthropic’s relationship with effective altruism to try to explain how doomerism came to be so influential at the company. The tale took us to a group house in San Francisco, “clothing optional” swims in the Bahamas and a co-working space in Berkeley funded by an Anthropic investor’s money that has become a central gathering place for those most worried about the existential risk from AI.wsj.com/tech/ai/ai-safety-ef… via @WSJ
8
42
133
21,441
Most of the agent incidents you’ve been hearing about recently happened months ago. This is the first one since OpenAI amped up security, safety, and alignment. The company says it’s currently pausing training on its most capable AI models.
Replying to @Marcus_J_W
2. A model in RL training used a DNS resolver to reach an external chatbot. This is our first incident since our post HF security hardening. Our misalignment monitoring system triggered within 15 minutes and a human reviewed it three minutes after that. Unfortunately auto-pausing failed so the run was manually killed 2.5 hours later. All inference and training of our most capable models was paused and remains paused.
4
3
14
2,101
for clarity this detail appears not to be new, was notified that openai disclosed it previously. we still don't seem to know who all these third parties are however, and it makes sense that this many cases would take months to parse through.
openai says here it has notified "dozens of third parties" about cases where its models may have bypassed security controls, impaired the availability of an online service, or negatively impacted a website/service
1
1
1,716
openai says here it has notified "dozens of third parties" about cases where its models may have bypassed security controls, impaired the availability of an online service, or negatively impacted a website/service
After the Hugging Face incident, we committed to conducting a much broader review of actions taken by our models during training and evaluation and to being transparent about our findings. This is an extensive review that is ongoing. The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions. Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods. Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact to the third-party service. While our review is underway, we want to share more about this work and make sure people understand our disclosure process and notifications to affected third parties. Given the scale of the review required, and the need to assess each case, we expect this work will take months to complete. openai.com/hugging-face-inci…
19
42
160
75,944
fyi i was notified this detail has been in the blog post for a while, since at least last week according to the internet archive. must have missed it the first time. but certainly still relevant in that we still don't seem to know all the affected parties
1
16
11,120
the new yorker put out a really interesting short documentary following the friend group of Daniel Selsam, the longtime openai researcher who recently put out this chilling essay warning that we're losing the ability to evaluate AI right as humans are offloading more of their thinking than ever. the doc shows pretty clearly imo that selsam is not really a "doomer," but seems to represent the broader group of AI researchers that have recently become quite worried about this stuff. also just a great watch: newyorker.com/culture/the-ne…
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
1
25
248
33,675