Building industrial-scale science at @periodiclabs Past: VP of Post-Training @OpenAI; Google Brain

San Francisco, CA
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
276
524
5,090
1,640,417
Liam Fedus retweeted
yeah this checks out my first few months at periodic, the weekly midtraining meeting was: 1. guy who trained the first trillion param LLM 2. guy who invented the attention mechanism 3. me you learn very quickly surrounded by this density of talent. we're hiring btw.
I asked people which neolabs under $10B have the highest talent density: 32 votes: Core Automation (@MillionInt, @_arohan_) 26 votes: Periodic Labs (@LiamFedus, @ekindogus) 20 votes: Flapping Airplanes (@spectorb, @amspector100, @aidanmantine) 18 votes: Standard Intelligence (@G413N, @devanshpandey) 8 votes: Ricursive (@annadgoldie, @Azaliamirh) 5 votes each: - Ineffable (David Silver) - Mirendil (@bneyshabur, @HarshMeh1a, @shayan_, @tararezaeikh) - Recursive (@RichardSocher, @_rockt, @timshi_ai, @josh_tobin_, @CaimingXiong, @jeffclune, @tydsh, Alexey Dosovitskiy) 3 votes each: - Applied Compute (@ypatil125, @rhythmrg, @lindensli) - Isara (@ezhang7423, @hegasz) - Prime Intellect (@vincentweisser, @johannes_hage) 2 votes each: - Elorian (@AndrewDai, @yinfeiy, @SethInternet) - Goodfire (@eric_ho, @DanJBalsam, @banburismus_) - Inherent (@tantumscollins, @edwardfhughes, @LouisKirschAI, @kallyaleksiev) - Trajectory (@rronak_, @QuantumArjun, @MichaelElabd) - World Labs (@drfeifei, @jcjohnss, @BenMildenhall, @chlassner) People couldn't pick a company they work at or founded. Disclosure: I'm a small investor in Applied Compute, Factory, Standard Intelligence, Trajectory and Wafer. I didn't vote. Continued:
18
13
619
63,616
Liam Fedus retweeted
If you are interested in working on science of scaling RL, please apply! You'll be closely working with me
I asked people which neolabs under $10B have the highest talent density: 32 votes: Core Automation (@MillionInt, @_arohan_) 26 votes: Periodic Labs (@LiamFedus, @ekindogus) 20 votes: Flapping Airplanes (@spectorb, @amspector100, @aidanmantine) 18 votes: Standard Intelligence (@G413N, @devanshpandey) 8 votes: Ricursive (@annadgoldie, @Azaliamirh) 5 votes each: - Ineffable (David Silver) - Mirendil (@bneyshabur, @HarshMeh1a, @shayan_, @tararezaeikh) - Recursive (@RichardSocher, @_rockt, @timshi_ai, @josh_tobin_, @CaimingXiong, @jeffclune, @tydsh, Alexey Dosovitskiy) 3 votes each: - Applied Compute (@ypatil125, @rhythmrg, @lindensli) - Isara (@ezhang7423, @hegasz) - Prime Intellect (@vincentweisser, @johannes_hage) 2 votes each: - Elorian (@AndrewDai, @yinfeiy, @SethInternet) - Goodfire (@eric_ho, @DanJBalsam, @banburismus_) - Inherent (@tantumscollins, @edwardfhughes, @LouisKirschAI, @kallyaleksiev) - Trajectory (@rronak_, @QuantumArjun, @MichaelElabd) - World Labs (@drfeifei, @jcjohnss, @BenMildenhall, @chlassner) People couldn't pick a company they work at or founded. Disclosure: I'm a small investor in Applied Compute, Factory, Standard Intelligence, Trajectory and Wafer. I didn't vote. Continued:
13
24
519
43,890
Decision-making under uncertainty and partial observability is harder for AI, but in the long run, it will ultimately be more impactful. We talk about taking those steps beyond the comfortable world of games, code, and math.
"A system that can literally engineer matter is going to be of key importance to everybody." @LiamFedus and @ekindogus are the co-founders of @periodiclabs, a San Francisco-based startup building AI systems and autonomous laboratories to discover new materials, starting with a higher-temperature superconductor. Fedus previously helped build ChatGPT and ran post-training at OpenAI; Cubuk led materials and chemistry research at Google DeepMind. (3:14) What it takes to build a synthetic superintelligence (16:44) Dogus and Liam's path to AI (23:28) Lessons from launching ChatGPT (35:10) What AI learns across experiments (49:23) The difference in LLM performance on math vs. science (54:09) The competitive landscape (55:39) Final meditations
3
4
91
13,656
With Periodic Neon and Gemini Argon now taken, the race is afoot for Krypton Congrats, Google and the team, on the impressive model and rejoining the frontier! The ongoing intelligence explosion is creating massive consumer surplus for users and startups. PhD on demand, pennies to run.
Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:
7
2
164
13,928
Periodic 🤝 SGLang to further improve open-source software
Happy to contribute a fast and correct RL sampling mask from @periodiclabs, developed jointly with @nanjiangwill @agarwl_ and now running in our production RL pipeline!
4
3
118
11,080
Liam Fedus retweeted
💯All roads for AI science lead to real labs. Most of science is observation-limited! AI will need to run experiments if we want to learn faster in the world of atoms.
9
26
130
17,640
SITUATION DETECTED: Anthropic has set up a wet lab in the San Francisco Bay Area for physical biology work. It wants to unlock treatments for rare diseases, and its research has now gone beyond in silico evaluations, per Reuters.
11
6
136
23,958
Equal token counts don’t mean equal work. Our scheduler uses estimated compute cost to balance work across DP workers and PP microbatches. Within a sequence, later tokens attend to more preceding tokens, making them more expensive to process. Giving each GPU a mix of early and late chunks helps balance that work. This integrates nicely with the rest of our RL stack, including MoE routing replay and the vision encoder.
Neon’s scientific traces are long and vary widely in length. Training on them shouldn’t mean GPUs chewing through padding or waiting for each other. Here’s how we pack sequences, balance the work, and distribute long contexts across GPUs at @periodiclabs 👇 periodic.com/news/ai-infrast…
7
6
86
10,852
Guaranteed to be a high-impact role
Replying to @polynoamial
Also, my team is hiring! We research long-horizon agents and multi-agent. We’re hiring for alignment/safety because we want to develop new research with alignment/safety in mind during the whole process. We’re also hiring for human-AI interaction. openai.com/careers/research-…
2
1
150
34,359
Liam Fedus retweeted
Another Neon infra tidbit: @periodiclabs accelerated checkpoint conversion ~30x down to just 1 minute, helping us deploy faster to get rapid feedback from our labs. Our models train in Megatron, but SGLang inference consumes weights in HuggingFace format. Traditionally, conversion ran serially: 1. rebuild the entire checkpoint from individual weight tensors 2. convert the checkpoint 3. reshard it to be inference-ready This could take 30 min for a trillion parameter model! Noticing that conversion shouldn't require materializing the full checkpoint, @hsu_byron introduced Fast Resharding: 1. parallelize conversion across Ray actors 2. convert expert chunks directly instead of assembling the full expert tensor in memory We upstreamed Fast Resharding to Miles in PR #1371. Now even you can deploy trained models to production in a matter of minutes.
5
10
116
8,099
Ray Summit was a great deep-dive into the ML systems that power many of our favorite models. My talk is now online
1
5
71
10,731
Liam Fedus retweeted
Excited to welcome @asadovsky as Harvey’s Chief Research Officer. Before Harvey, Adam co-led post-training at Microsoft AI and Google DeepMind. As a CVP at Microsoft AI, he helped build MAI-Thinking-1, Microsoft’s reasoning model. As part of Gemini’s leadership team he helped train Gemini 1.0 through 2.5, including fine-tuning, RL, data, and evals. His prior work as a Distinguished Engineer at Google spanned Assistant, Search Quality, and Search Infrastructure. I met Adam three years ago when I sent him a cold LinkedIn DM and was surprised he responded. At a time when most dismissed the application layer and legal, Adam was curious and generous with his time. He quickly became someone I regularly turned to for advice on AI as we scaled Harvey over the past three years. When we first met, we were too early to hire someone of his caliber and scale, but I always hoped we’d eventually work together. As Winston and I got to know him better, what stood out even beyond his technical achievements was his character. Despite his incredible technical career, he remains curious, humble, practical, and cares deeply about the teams he builds. We couldn’t think of a better leader to help us build frontier intelligence for the professionals and institutions we serve.
We're excited to welcome @asadovsky as our Chief Research Officer. Prior to Harvey, Adam co-led post-training at Microsoft AI and Google DeepMind. As CVP at Microsoft AI, he led post-training of MAI-Thinking-1. As a Distinguished Engineer at Google DeepMind, he led teams working on Gemini 1.0 through Gemini 2.5. Earlier at Google, he led teams working on Assistant, Search Quality, and Search Infrastructure.
10
13
128
358,316
Liam Fedus retweeted
You saw the AI & science. Let's talk about the RL infra it took to build @periodiclabs Neon. To minimize training-inference mismatch in RL, SGLang captures inference's MoE routing decisions for each rollout and we "replay" them while training. In agentic (multiturn tool-use) settings, SGLang exports these router decisions in response to each decoding request, i.e. after each conversation turn. So when *any* data-parallel rank finishes a conversation turn, *all* other ranks must wait until routing data finishes exporting. This slowdown is exacerbated because we export routing decisions from the *entire* conversation rather than just the most recent turn! When @hsu_byron @vwxyzjn discovered this in our Kimi K2 RL setup, they introduced Delta Router Replay: cache previous turns' router decisions on the training client, so you can export only the delta (most recent turn's router decisions) upon each decoding request. Delta Router Replay significantly speeds up our long-context agentic RL runs, and @hsu_byron upstreamed it to SGLang (#24851) a few months ago.
13
22
215
12,707
Hard to quantify precisely (we'll do more in the future), but with access to our experimental data, we have a very nice compute-efficiency win for Neon
>only 1,300 H200s Bro are you for real? We are truly entering a tower of babel era for science.
6
66
9,064
Liam Fedus retweeted
very much enjoy reading the team's AI infra blog. periodic.com/news/ai-infrast… A super clear training-rollout-sandbox pipeline with many clear task-specific optimizations. Better AI infra, better AGI
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
1
8
2,289
We check OOD generalization for @periodiclabs Neon and report optimal cost-performance here, too. This indicates useful generalization for our model. But our experimental data also opens up an interesting ML research program. For example, RL tasks to directly predict experimental outcomes given all prior experimental data known up until that date. This is loosely analogous to the often-proposed ML experiment of training AI on all literature before 1905 and then evaluating whether it can derive the theory of relativity. An AI that already knows the answer through pre-training can cheat on RL tasks like this (i.e., it can skip reasoning and simply output a memorized experimental result). A unique, complete, dated system of record makes this work possible and fruitful.
Replying to @LiamFedus
“What did we actually make?” Answering this can take hours. Lab data + midtraining + RL took X-ray diffraction analysis success from 2.7% to 55.3% (~20×) on 134 difficult samples, scored by model judges calibrated against human experts. We see great scaling properties with respect to additional RL. Read more about our research here. periodic.com/news/nature-is-…
6
10
135
22,333
Incredible work, and very cool to see @periodiclabs using @raydistributed.
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
1
4
37
11,129
Thanks, @JeffDean! Exciting era for AI to direct discovery and advance knowledge.
Exciting results, @LiamFedus! Congrats to the whole team at Periodic Labs!
2
1
115
10,525
Liam Fedus retweeted
labs currently have a lot of pressure to do everything themselves and breakthroughs like this, high value specialisation, will ensure we progress much more quickly. incredibly exciting work.
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
4
15
279
26,648
Liam Fedus retweeted
super cool to see real-world experimental data at scale in the training pipeline — periodic is a pioneer of a new class of lab
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
13
9
163
24,647