Frontier AI Data Lab advancing AI through better data

Redwood City, California
We raised $350M at a $3.5B valuation, co-led by @insightpartners and @S32_VC with significant participation from existing investor Addition, as we surpassed $375M in ARRR, growing 18x+ in less than a year. The bottleneck to AI progress is no longer just models. It’s data. We’re building the frontier data lab to solve it. The round included participation from @AllegisCapital, @blumbergcapital, @MarchCPs, Factory, @Frontlinevc, @GreylockVC, @GVteam, @lightspeedvp, Prosperity7 Ventures, @Standardvc, @ThirdPointLLC, @walden_catalyst and @WellsFargo.
6
23
94
8,442
The bar for trustworthy, lasting benchmarks is rising. @StevenDillmann, Russell Yang, @GOrlanski, and @vincentsunnchen discuss what it takes to build great open benchmarks at Frontier Data Summit. frontierdatasummit.ai/
1
4
17
790
Snorkel AI retweeted
🎉 Week 132 of Venture with Grace Another packed week of conversations spanning AI inference, fintech, climate, the future of hiring, data-centric AI, agent trust, robotics, and physical AI. Week 132 lineup: @GeorgeHuSF, @FireworksAI_HQ President on AI Inference and Enterprise How AI inference is evolving and what it means for enterprise adoption. Fabiola Quinzanos, Monashees Partner on Fintech, Climate and Consumer Investing across fintech, climate, and consumer technology. @hkolam, @FindemAI CEO on AI Agents and the Future of Hiring How AI agents are transforming recruiting and talent management. @ajratner, @SnorkelAI CEO on Data-Centric AI and AI Development Building better AI systems through data-centric development. @BarrMoses_MC, @montecarlodata CEO on Agent Trust, AI Infrastructure and Data Building trust and reliable infrastructure for the agentic AI era. Ryan Gibson, @EclipseVentures on Physical AI, Robotics and Industry Where physical AI and robotics are creating the next wave of industrial innovation. Another strong week of conversations on AI infrastructure, investing, enterprise adoption, talent, data, and physical AI, highlighting how quickly AI is moving from software into every layer of the economy. More coming next week. See you live 👋 #VentureWithGrace #AI #Startups #Founders #VC
2
1
8
390
Snorkel AI retweeted
Releasing PhantomEnvironments today! Fully synthetic RL environments that turn off-the-shelf LLMs into strong search agents: a 7B LLM trained with RL in PhantomEnvs performs like an agent 10x its size. You can generate environments at $0 cost, free of distillation, and fully open-source. And we are just getting started: we won a large GPU compute grant from @NVIDIAAI to continue our work on synthetic data and self-improving agents. We'll be "on tour" in SF next week: - @SnorkelAI Frontier Data Summit (Oct 8) - Oral talk at the @COLM_conf LSEI workshop and poster at LLA workshop (Oct 9) Drop me a message if you'd like to chat! Paper: arxiv.org/abs/2609.40221 Code: github.com/kilian-group/phan… (coming soon; we're finishing a hero run on our new NVIDIA compute) --- Agent training is bottlenecked by RL environments. Our bet: the biggest lever is data, and the way to scale it is synthetic environments. Earlier this year we asked a simple question: can an LLM learn the generalizable skill of agentic search (decomposing a question, retrieving documents, composing knowledge) from environment interaction alone? And to push it to the extreme: can it learn this in worlds where the questions are templated ("Who is the mother of Alice's father?") and the Wikipedia is generated by rules? Yes. Agents learn the skill and transfer to real-world search remarkably well. PhantomEnvs lets you scale that data axis in agent training for free. PhantomEnvs is a milestone in my PhD, and the latest in a line of Phantom* projects. When I started the PhD two years ago, OpenAI o1 had just come out and Claude 3.x was getting its first tool-use abilities. I became obsessed with training these systems myself, squeezing every GPU I could find to understand what goes into building something so cool. I thought it would take 5 years to get a grip, but it's been 2. - PhantomWiki: evaluating LLM reasoning in isolation from memorization - PhantomReasoning: showing LLMs can be post-trained on synthetic data - PhantomEnvironments: training LLM search agents on rule-generated synthetic RL environments There's always another Phantom* project on the way... much to my lab's annoyance because of the number of phantom-* slack channels. RL and synthetic data are two deep gold mines, and we (are starting to) understand how to dig. Joint work with Swathi Saravana Selvam, @GongAlbert, @ChaoWan0331, Christian Belardi, Dongyoung Go, @katielulula, and @KilianQW. Big thanks to @NVIDIAAI for the continued support!
15
14
99
4,875
Snorkel AI retweeted
Thanks for the great conference + session @modal ! So much whitespace to cover on both public & private internal evals, as our ability to rigorously *evaluate* AI lags its advancement rate for the first time in AI history
The line to attend the evals track at Modal’s Runtime conference was practically out the door. This is why! @ajratner
2
30
2,377
One week 👀 frontierdatasummit.ai
1
2
9
768
Our research team has five papers accepted to @NeurIPSConf 2026: - Agents' Last Exam: long-horizon professional tasks, <1% average pass rate on the hardest tier (arxiv.org/abs/2606.05405) - Continual Learning Bench: do agents improve with experience? (arxiv.org/abs/2606.05661) - JudgmentBench: expert attorney preference judgments for ranking legal AI output (arxiv.org/abs/2605.25240) - SkillOrchestra: routing tasks by agent competence and cost (arxiv.org/abs/2602.19672) - SlopCodeBench: agent-written code gets more bloated and eroded with every iteration (arxiv.org/abs/2603.24755) Congrats to @amanda_dsouza, @vincentsunnchen, @chris_m_glaze, @RamyaRamakri, @fredsala, Charles Dickens, and to all our collaborators!
3
14
74
4,943
TLAPS-Bench uses the TLA+ Proof System to formally prove the correctness of complex protocols and systems. It also provides a rich, extensible, and scalable benchmark suite for frontier AI. From @tianyin_xu, one of 25+ posters at Frontier Data Summit. frontierdatasummit.ai/
3
8
2,672
Snorkel AI retweeted
The next frontier lab featuring Terminal-Bench-Science, with a big performance jump from 12.4% (Gemini 3.8 Flash) to 57.6% by @GoogleDeepMind's brand-new Gemini 4 Argon - #3 on our leaderboard. If you're a scientist, I'd be curious to hear about your experience using this model over the coming weeks!
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
4
5
38
3,528
Snorkel AI retweeted
Excited to see @GOrlanski's new benchmark LibraryDesignBench measuring how agents can build code libraries for other agents! And excited to support via @SnorkelAI Open Benchmarks Grants
Agents will soon replace humans as the main users and designers of libraries. A good library will let future agents write correct programs with less code. Introducing LibraryDesignBench: one agent designs a library, and other agents write programs with it.
1
11
1,010
Can agents design libraries other agents can use? LibraryDesignBench tests whether libraries built by one agent help others write simpler, correct code. @GOrlanski led the project during his time as a Snorkel AI research fellow, with support from Snorkel AI's Open Benchmarks Grants program.
Agents will soon replace humans as the main users and designers of libraries. A good library will let future agents write correct programs with less code. Introducing LibraryDesignBench: one agent designs a library, and other agents write programs with it.
3
18
1,034
coding agents are becoming the primary producers of software, and increasingly its primary users too. how good are they at writing libraries for other agents? we introduce library design bench, led by @Gorlanski: an agent is given a deliberately vague spec and writes a library. we then score that library only by how other agents fare using it (e.g. correctness & simplicity of downstream code) we find that opus 5.5 already designs libraries that help downstream agents more than the human-written ones do (same pass rate, less code). but there's still headroom! 82% of issues with verbosity / rigidity issues in downstream code trace back to the library itself!
Agents will soon replace humans as the main users and designers of libraries. A good library will let future agents write correct programs with less code. Introducing LibraryDesignBench: one agent designs a library, and other agents write programs with it.
2
16
654
Snorkel AI retweeted
Agents coding for agents---where are we at? Excited to share LibraryDesignBench, a new benchmark designed to measure exactly this question. ldbench.com/, exciting new work from @GOrlanski
Agents will soon replace humans as the main users and designers of libraries. A good library will let future agents write correct programs with less code. Introducing LibraryDesignBench: one agent designs a library, and other agents write programs with it.
1
6
29
1,012
Snorkel AI retweeted
This was an exciting project! It would not be possible without my amazing collaborators: @a1zhang @atrost3122 @vincentsunnchen @fredsala @awsTO @lschmidt3 This work was generously supported by @DARPA, @NSF, @PrimeIntellect, and through @SnorkelAI 's Open Benchmark Grant.
1
1
17
515
Snorkel AI retweeted
Agents will soon replace humans as the main users and designers of libraries. A good library will let future agents write correct programs with less code. Introducing LibraryDesignBench: one agent designs a library, and other agents write programs with it.
14
27
128
17,383
We will be presenting FutureSim at the Frontier Data Summit and also in COLM workshops next week in SF! Come join us on October 8th alongside other frontier benchmarks by registering here: frontierdatasummit.ai/
FutureSim is a live backtesting simulation testing how models adapt at test time to new information and forecast outcomes. Runs span 90 simulated days, 5,000+ tool calls, and 12+ hours per model. From @ShashwatGoel7, @nikhilchandak29, and @arvindh__a. One of 25+ posters at Frontier Data Summit frontierdatasummit.ai/
3
3
33
3,807
Snorkel AI retweeted
An agent should become more useful the longer it knows you. Not slower or more expensive. Hear Co-founder @sudip_r0y speak about agent memory alongside @nikogrupen at @harvey at @SnorkelAI's Frontier Data Summit on October 8th.
4
9
32
2,291
ORCA-Bench puts general-purpose coding agents in a production-fidelity on-call setting. Agents tackle 1,079 realistic root cause analysis tasks using six days of metrics, logs, traces, and source code. From Kyuseong Choi and Albert Gong, along with @traversal_ai; one of 25+ posters at Frontier Data Summit. frontierdatasummit.ai/
1
1
15
711
awesome discussion with @jerryjliu0 and a great group! one theme: data must be developed to be useful users/enterprises have a trove of interesting & valuable raw data (e.g. patient records, historical images/documents), but making that ready for evaluation and training requires a new form of engineering & research. this requires new interfaces & programming models involving subject matter experts to make the resulting data & environments sufficiently difficult/fair/complex!
Yesterday I hosted a fun dinner conversation with @vincentsunnchen from @SnorkelAI on evals and RL environments. The "data and RL env" companies (like Snorkel) have seen massive growth in the past few years. There's been an explosion of interest in evals. At the same time, models are ripping through benchmarks with each new release. We talked about the evals everyone is defining, what evals are still left unsolved, what’s left up to frontier models vs. intelligence that you own, and more: - A big challenge for building RL environments is “fairness” - when the model fails on a given environment, can you attribute it to the input, harness, or reward model? - Building proper rewards is hard. Some tasks are not easily quantifiable. You also want to discourage reward hacking. At the same time, you don’t want to be too prescriptive with intermediate rewards. - Long horizon evals are still extremely hard, some business processes can take up to weeks or months before the final outcome - Most regulated industries still need human in the loop to guarantee ~100% accuracy, “80%” accuracy is not good enough - As models get more intelligent, there will be a barbell of boutique data vendors (e.g. any SMB) any scaled up data providers. - Models still exhibit “jagged intelligence” where they still fail on a long tail of edge cases. - There might always be opportunities to gather unique data for a given task and posttrain models for lower cost and higher accuracy. This marks #003 in our founder dinner series. What topic should we discuss next? Let us know your thoughts below!
4
17
1,920
Snorkel AI retweeted
Yesterday I hosted a fun dinner conversation with @vincentsunnchen from @SnorkelAI on evals and RL environments. The "data and RL env" companies (like Snorkel) have seen massive growth in the past few years. There's been an explosion of interest in evals. At the same time, models are ripping through benchmarks with each new release. We talked about the evals everyone is defining, what evals are still left unsolved, what’s left up to frontier models vs. intelligence that you own, and more: - A big challenge for building RL environments is “fairness” - when the model fails on a given environment, can you attribute it to the input, harness, or reward model? - Building proper rewards is hard. Some tasks are not easily quantifiable. You also want to discourage reward hacking. At the same time, you don’t want to be too prescriptive with intermediate rewards. - Long horizon evals are still extremely hard, some business processes can take up to weeks or months before the final outcome - Most regulated industries still need human in the loop to guarantee ~100% accuracy, “80%” accuracy is not good enough - As models get more intelligent, there will be a barbell of boutique data vendors (e.g. any SMB) any scaled up data providers. - Models still exhibit “jagged intelligence” where they still fail on a long tail of edge cases. - There might always be opportunities to gather unique data for a given task and posttrain models for lower cost and higher accuracy. This marks #003 in our founder dinner series. What topic should we discuss next? Let us know your thoughts below!
13
6
31
5,395
Reimagining a Science of AI Measurement & Evaluation @sanmikoyejo of @stai_research and @VirtueAI_co sits down with @alexgshaw, co-creator of @terminalbench, for a conversation at Frontier Data Summit frontierdatasummit.ai/
6
18
876