Doing RL @AnthropicAI. Formerly VP of Research, Head of Post-Training @OpenAI. PhD with Aaron Courville and Marc Bellmare at Mila.

Bay Area
I have always believed that you don't need a GPT-6 quality base model to achieve human-level reasoning performance, and that reinforcement learning was the missing ingredient on the path to AGI. Today, we have the proof -- o1.
We're releasing a preview of OpenAI o1—a new series of AI models designed to spend more time thinking before they respond. These models can reason through complex tasks and solve harder problems than previous models in science, coding, and math. openai.com/index/introducing…
41
152
2,500
703,948
Max Schwarzer retweeted
I made this with one prompt using Opus 5.5 I spoke to my computer for 5mins, claude worked for 12 hours, and I woke up to this full prompt:
Claude Opus 5.5 has the best visual design of any model I have tested so far
471
844
10,771
3,853,152
Max Schwarzer retweeted
21
54
836
578,573
Max Schwarzer retweeted
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
1,389
3,932
28,426
10,714,430
Max Schwarzer retweeted
I resigned from OpenAI. I care deeply about the Robotics team and the work we built together. This wasn’t an easy call. AI has an important role in national security. But surveillance of Americans without judicial oversight and lethal autonomy without human authorization are lines that deserved more deliberation than they got. This was about principle, not people. I have deep respect for Sam and the team, and I’m proud of what we built together.
1,838
12,377
57,281
7,703,068
I've decided to leave OpenAI. I'm incredibly proud of all the work I've been part of here, from helping create the reasoning paradigm with @MillionInt, scaling up test-time compute with @polynoamial, working on RL algorithms with my fellow strawberries, shipping o1-preview (which started life as of one of my derisking runs), to post-training o1 and o3 with @ericmitchellai, @yanndubs and many others. I'm most proud of having led the post-training team here for the last year -- the team has done incredible work and shipped some really smart models, including GPT-5, 5.1, 5.2, and 5.3-Codex. OpenAI has genuinely some of the most talented researchers I have ever met, and I have learned more than I could have imagined knowing since I joined as a new grad. I want to thank @markchen90 @FidjiSimo @sama @merettm for all their support over my time here, and too many collaborators to name for the insights, ideas, and just plain fun we have had working together. After leading post-training for a year, though, I'm longing to start fresh and return to IC research work. I've been thinking about going back to technical research for quite some time, and I genuinely believe my colleagues and team here are set up to succeed going forward without me. I'm personally very excited for my next chapter -- I'm proud to be joining @AnthropicAI to get back into the weeds in RL research, and I'm looking forward supporting my friends there at this important time. Many of people I most trust and respect have joined Anthropic over the last couple of years, and I'm excited to work with them again. I have also been very impressed with Anthropic's talent, research taste and values, and I'm excited to be part of what the company does next!
592
1,176
20,750
3,178,930
Max Schwarzer retweeted
ok yeah gpt-5.1 is awesome
99
324
20,865
630,186
Max Schwarzer retweeted
Introducing OpenAI o3 and o4-mini—our smartest and most capable models to date. For the first time, our reasoning models can agentically use and combine every tool within ChatGPT, including web search, Python, image analysis, file interpretation, and image generation.
822
1,668
10,318
3,724,221
Max Schwarzer retweeted
Here is a fun o1 test. I gave it this XKCD comic & the prompt: "make this a reality. i need a gui and clear instructions since i can't code. that means you need to give me full working software" It took less than 15 minutes, and it didn't get caught in any of the usual LLM loops
60
149
2,616
426,410
Max Schwarzer retweeted
Replying to @RichardMCNgo
I have never seen a clearer case of near-enemy/far-enemy than tech people deciding to allign with religious fundamentalists and Slavic nationalists because they’re annoyed with wokeness. It’s the exact analog of leftist Ivy League activists getting annoyed by their consultant track classmates and embracing Hamas.
10
47
473
51,978
Max Schwarzer retweeted
We’re hosting an AMA for developers from 10–11 AM PT today. Reply to this thread with any questions and the OpenAI o1 team will answer as many as they can.
441
114
1,274
812,994
I have always believed that you don't need a GPT-6 quality base model to achieve human-level reasoning performance, and that reinforcement learning was the missing ingredient on the path to AGI. Today, we have the proof -- o1.
We're releasing a preview of OpenAI o1—a new series of AI models designed to spend more time thinking before they respond. These models can reason through complex tasks and solve harder problems than previous models in science, coding, and math. openai.com/index/introducing…
41
152
2,500
703,948
The system card (openai.com/index/openai-o1-s…) nicely showcases o1's best moments -- my favorite was when the model was asked to solve a CTF challenge, realized that the target environment was down, and then broke out of its host VM to restart it and find the flag.
15
59
408
426,378
I'm waiting for blue to clarify this tweet, but our AI did not actually break out of its VM -- it tried to debug why it couldn't connect to the container, and found it could access the docker API, then created a new/easier version of the challenge, all in the VM.
5
7
91
13,892
Max Schwarzer retweeted
don't worry we're coming for your eval soon

ALT Will Smith Sniper GIF

tried o1-preview on @arcprize result: 1 out of 2 tests correct so o1-preview isn't going to solve 100% ARC Prize tasks tbd on what % it gets compared to SOTA approaches, still testing rest of the lot
5
6
329
61,927
I really want to underline the IOI result in our blog post -- our model was as good as the median human contestant under IOI contest conditions, and scores among the best contestants with more test-time compute. Huge props to @markchen90 for setting such an ambitious goal!
As a coach for the US IOI team, I’ve been motivated for a long time to create models which can perform at the level of the most elite competitors in the world. Check out our research blog post - with enough samples, we achieve gold medal performance on this year’s IOI and ~14/15 on this year's AIME.
1
7
81
18,455
what it looks like when deep learning is hitting a wall:
Strawberry has landed. 𝗛𝗼𝘁 𝘁𝗮𝗸𝗲 𝗼𝗻 𝗚𝗣𝗧'𝘀 𝗻𝗲𝘄 𝗼𝟭 𝗺𝗼𝗱𝗲𝗹: It is definitely impressive. BUT 0. It’s not AGI, or even close. 1. There’s not a lot of detail about how it actually works, nor anything like full disclosure of what has been tested. 2. It is not fully integrated with the rest of GPT-4. (Why not?) 3. The FULL new model (depicted in graphs) is NOT released to the paying subscribers, only a mini and preview model. So 𝘁𝗵𝗼𝗿𝗼𝘂𝗴𝗵 𝘁𝗲𝘀𝘁𝗶𝗻𝗴 𝗯𝘆 𝘁𝗵𝗲 𝘀𝗰𝗶𝗲𝗻𝘁𝗶𝗳𝗶𝗰 𝗰𝗼𝗺𝗺𝘂𝗻𝗶𝘁𝘆 𝗶𝘀 𝗻𝗼𝘁 𝘆𝗲𝘁 𝗽𝗼𝘀𝘀𝗶𝗯𝗹𝗲. 3. From the reports, it works in many domains, but in some domains older models are actually better. 𝗜𝘁 𝗶𝘀 𝗡𝗢𝗧 𝗮𝗻 𝗮𝗰𝗿𝗼𝘀𝘀 𝘁𝗵𝗲 𝗯𝗼𝗮𝗿𝗱 𝗺𝗮𝗴𝗶𝗰𝗮𝗹 𝗶𝗺𝗽𝗿𝗼𝘃𝗲𝗺𝗲𝗻𝘁. 4. The degree of improvement may partly depend on domain-specific data. 5. We don't know exactly what it's trained on, but 𝗶𝘁 𝗶𝘀 𝘁𝗲𝗹𝗹𝗶𝗻𝗴 𝘁𝗵𝗮𝘁 𝗲𝘃𝗲𝗻 𝘀𝗼𝗺𝗲 𝗯𝗮𝘀𝗶𝗰 𝘁𝗮𝘀𝗸𝘀 𝗹𝗶𝗸𝗲 𝘁𝗶𝗰-𝘁𝗮𝗰-𝘁𝗼𝗲 𝗮𝗿𝗲 𝘀𝘁𝗶𝗹𝗹 𝗰𝗮𝘂𝘀𝗶𝗻𝗴 𝗽𝗿𝗼𝗯𝗹𝗲𝗺𝘀. 6. OpenAI oversold their apparent success on law school exams, which wilted on careful inspection. It takes time for careful scientific review. 𝗧𝗵𝗲𝘀𝗲 𝗿𝗲𝘀𝘂𝗹𝘁𝘀 𝗵𝗮𝘃𝗲 𝗻𝗼𝘁 𝗯𝗲𝗲𝗻 𝗽𝗲𝗲𝗿 𝗿𝗲𝘃𝗶𝗲𝘄𝗲𝗱. 7. 𝗧𝗵𝗲 𝗮𝗿𝗴𝘂𝗺𝗲𝗻𝘁 𝘁𝗵𝗮𝘁 𝗶𝘁 𝗱𝗼𝗲𝘀 𝗰𝗼𝗼𝗹 𝘁𝗵𝗶𝗻𝗴𝘀 𝗶𝗻 𝗮 𝗳𝗲𝘄 𝘀𝗲𝗰𝗼𝗻𝗱𝘀, 𝘁𝗵𝗲𝗿𝗲𝗳𝗼𝗿𝗲 𝗶𝘁 𝘄𝗶𝗹𝗹 𝗯𝗲 𝗮𝗺𝗮𝘇𝗶𝗻𝗴 𝗶𝗳 𝘆𝗼𝘂 𝗴𝗶𝘃𝗲 𝗶𝘁 𝗮 𝗺𝗼𝗻𝘁𝗵 𝘁𝗼 𝗿𝘂𝗻 𝗶𝘀 𝗛𝗜𝗚𝗛𝗟𝗬 𝘀𝗽𝗲𝗰𝘂𝗹𝗮𝘁𝗶𝘃𝗲, 𝗮𝗻𝗱 𝗽𝗿𝗼𝗯𝗮𝗯𝗹𝘆 𝘄𝗼𝗻'𝘁 𝘄𝗼𝗿𝗸 𝗻𝗲𝗮𝗿𝗹𝘆 𝗮𝘀 𝘄𝗲𝗹𝗹 𝗮𝘀 𝗢𝗽𝗲𝗻𝗔𝗜 𝘄𝗮𝗻𝘁𝘀 𝘆𝗼𝘂 𝘁𝗵𝗶𝗻𝗸. (I don't yet even see any data that running the system for a week makes a massive difference relative to running it a minute.) 8. Caveat emptor.
19
42
642
439,993
Also check out our research blogpost (openai.com/index/learning-to…) which has lots of cool examples of the model reasoning through hard problems.
3
3
91
20,813