Researching reasoning @OpenAI | Co-created Libratus/Pluribus superhuman poker AIs, CICERO Diplomacy AI, and OpenAI o-series 🍓 reasoning models

San Francisco, CA
I met @ssokota 5 years ago at a conference poster session. I was impressed with his explanation, so I offered him an internship. We continued collaborating and he now works with me at @OpenAI. Later, he told me I was the only person to stop by his poster for the entire session.
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
32
35
1,143
95,983
Replying to @longphann_
@polynoamial Results are even more impressive at max! 6.1 sol outperforms Opus 5.5 while using 4x fewer tokens.
1
4
64
7,518
I’m happy that @OpenAI presents model evals this way. We should measure intelligence as a function of cost.
GPT-6.1 Sol: near-Astra intelligence for a fifth of the price. It’s the most cost-efficient model for its performance available today.
80
59
1,248
86,707
Was testing out dots over the weekend and it figured out how to save me ~$500/yr in recurring charges. Most impressive part was when it connected to customer service and handled a whole texting convo on my behalf.
Introducing dots, powered by GPT-6 Astra. Remarkably capable, always-on agents built to handle everything.
49
27
702
83,981
GPT-6 Sol and Luna are out, and they are better AND 50% cheaper than 5.6. Luna is now $0.10 input / $0.50 output per 1M tokens. This is on top of the 80% price cut to Luna we made at the end of July. Output went from $6 -> $0.50 within two months.
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
57
70
1,421
126,851
A few thoughts on this: 1) If you’ve only seen clips of this interview, I’d encourage you to watch the full podcast. I push back on plenty of AI hype in it. 2) As I said in the podcast, this example is academic. My intention was to illustrate how hard it is to make absolute guarantees about isolation, which is why it's important to have layers of defense. The part before the clip starts is me talking about other layers of defense. 3) The example I'm bringing up isn't about weight exfiltration via temperature sensors, it's about coordination between agents that are supposed to be fully isolated and independent. Coordination can require very few bits of information. 4) One lesson from the HF incident is that we put too much trust in sandbox isolation and didn't have enough independent safeguards. Airgapping is an extremely strong safeguard. When designing safety protocols, I think it's much better to overestimate rather than underestimate.
OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: firesidealpha.substack.com/p…
197
140
1,764
294,135
Full episode:
New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
5
3
82
18,913
Happy to finally do a deep dive on multi-agent with @dwarkesh_sp! None of it would have happened without the great work on multi-agent from my @OpenAI teammates @kevinleestone, @mikegmalek, @__eknight__, @amuellerml, @zhangir_azerbay, @CheukHeiChu, and many others.
New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
51
100
1,308
134,280
Also, my team is hiring! We research long-horizon agents and multi-agent. We’re hiring for alignment/safety because we want to develop new research with alignment/safety in mind during the whole process. We’re also hiring for human-AI interaction. openai.com/careers/research-…
17
22
674
115,019
Noam Brown retweeted
The Navier Stokes solution was the result of a collaboration of ~10,000 (!) agents working together. Over the past year, we’ve been training models to collaborate through multiagent RL. It’s been amazing to see how much better the models have become at this: it seems clear now that one of the most effective ways to solve hard problems is to give models huge amounts of unstructured parallel test-time compute and let them decide how to work together.
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
19
35
384
109,163
Noam Brown retweeted
One of the most amazing moments for me in OpenAI history was watching this happen over the past week:
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
1,084
938
17,236
1,677,113
Very sad to see Levent double down on the plagiarism accusation. I hope my friends at @AnthropicAI stand up to this internally. It should be clear by now what the truth is.
152
54
2,495
317,226
Happy to see some @AnthropicAI employees are willing to speak out at least
fwiw I think it is _extremely_ unlikely that user data had any influence here - there is no way OAI would pull user transcripts for this, or knowingly train on it in a way that would've influenced this. I think its pretty important people don't run away with 'your user data isn't safe in codex' - because it surely is (based on everything I can assume from the outside)
18
6
435
87,235
RT @deanwball: “To solve the Navier-Stokes problem, we used an internal model that is significantly more capable than GPT-6 Astra.” https:/…
33
1,410
Noam Brown retweeted
another crazy day working at the crazy day factory
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
4
7
174
8,320
Seeing this new internal model solve open after open math problem shortly after training commenced was the wildest thing I have ever witnessed at my time at OpenAI
Replying to @OpenAI
This model represents a step-function improvement on many benchmarks, and its training is ongoing. Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the strict safeguards—including monitoring and isolation—that we apply to all our frontier evaluations.
22
65
1,123
133,359
Noam Brown retweeted
This past week at OpenAI has been the most humbling experience of my life. We all feel the weight of what’s coming. When I joined OpenAI earlier this year, I speculated that AI might write an Annals paper in 2027—and maybe, just maybe, solve a Millennium Prize Problem in 2028. That was considered very bullish at the time. The first happened in May; the second, this September. What a surreal time to be alive!
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
51
101
1,316
110,560
Noam Brown retweeted
Last July, when LLMs won gold medals at the IMO, I felt that someday AI might be able to solve some of the hardest problems in math. Now, an internal OpenAI model (still improving!) has resolved a Millennium Prize Problem. It feels surreal that that day came so, so soon!
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
14
38
532
48,015