@SnorkelAI @uwcse / prev @StanfordAILab – Interested in data management systems for machine learning, weak supervision, and impactful applications.

Menlo Park, CA
I'm excited to announce @SnorkelAI's $350M Series E at $3.5B, led by @insightpartners and @S32_VC. We've grown 18x+ in the last 12 months since launching our Data-as-a-Service offering, passing $375M ARRR this week. As AI advances to superhuman capabilities, AI data & environment development must advance with it - and basic staffing and crowdsourcing approaches are not enough. AI progress now requires deep research and technology work that combines human expertise with specialized AI in compounding ways. @SnorkelAI is building the RSI data engine and frontier data lab for this next phase. We're honored to have the support of existing investors Addition, @lightspeedvp, @GreylockVC, @GVteam, P7, Factory, @WellsFargo, Walden Catalyst Ventures, and new investors @ThirdPointLLC, @MarchCPs, @BlumbergCapital, @AllegisCapital, @Frontlinevc, and @standard_vc. – @SnorkelAI started as a research project a decade ago at @StanfordAILab. Our thesis was simple: AI progress would become increasingly data-centric – and therefore data development should be studied as a true research and technology problem, not just a staffing and crowdsourcing one. Today, as AI capabilities verge on superhuman, building the data and environments to safely measure and train AI is becoming too hard for even the smartest human experts to do alone. Only humans and AI agents, collaborating together in compounding ways, can meet the accelerating needs of the frontier, and keep humans in the driver’s seat of AI progress for decades to come. At @SnorkelAI, we are building the data lab to define the shape of this new “Data 2.0” frontier, and the new paradigms of human-computer interaction needed to advance it. Our key focus is building the RSI engine for data, where specialized AI models accelerate and improve human expert output, and in turn, scaled human supervision is used to continuously evaluate and improve these models – creating a powerful compounding loop to keep pace with an accelerating RSI frontier. With this round of funding, we are also doubling down on our commitments to support data development for open benchmarking and evaluation (more news here soon!); an increasingly diverse ecosystem of general and specialized intelligence; and a path to safe, well-aligned AI built on robust training and evaluation data. Data development will guide and drive the next stages of AI – and must do so in a human-centric, AI accelerated, open, diverse, and safe way. We are excited to support this mission in the next decade of research ahead at @SnorkelAI. More thoughts here: snorkel.ai/blog/data-2-0-and…
73
115
412
89,311
Alex Ratner retweeted
I made lasteval.com. The premise is simple: AI is already improving lives, and we should have a page that collects these wins. We only talk about how "intelligent" the models are, what jobs they might automate, the trillions being invested, who will "win" the AI race, and the safety risks. All these things naturally worry people, and make them averse to AI. I think we're missing the point. The goal of technology should be to improve people's lives. AGI, ASI, RSI are all just in-group goals, that aren't obviously beneficial on their own. Speculative promises shouldn't be all we have to offer. Imagine releasing a new model and saying "it will score 99% on these benchmarks one day" instead of reporting what it already does today. I've spent the last few weekends reading about AI systems, both generative and predictive, that are already making a large positive impact on the world. It gives me hope, and I made this page so it does the same for others. Some examples in 🧵 Ultimately, the final, and only evaluation for AI that matters is whether it improved lives. The models are already so good, and yet the world is not that different. I think it's important to change that, and this is my attempt to paint the target. There's no way the list is complete, so if you think an AI system has had a large positive impact, do share it through the submit page!
7
14
82
5,603
Very proud of @SnorkelAI 's research work & contributions to some great projects heading to @NeurIPSConf 2026!
Our research team has five papers accepted to @NeurIPSConf 2026: - Agents' Last Exam: long-horizon professional tasks, <1% average pass rate on the hardest tier (arxiv.org/abs/2606.05405) - Continual Learning Bench: do agents improve with experience? (arxiv.org/abs/2606.05661) - JudgmentBench: expert attorney preference judgments for ranking legal AI output (arxiv.org/abs/2605.25240) - SkillOrchestra: routing tasks by agent competence and cost (arxiv.org/abs/2602.19672) - SlopCodeBench: agent-written code gets more bloated and eroded with every iteration (arxiv.org/abs/2603.24755) Congrats to @amanda_dsouza, @vincentsunnchen, @chris_m_glaze, @RamyaRamakri, @fredsala, Charles Dickens, and to all our collaborators!
24
1,562
Alex Ratner retweeted
Our research team has five papers accepted to @NeurIPSConf 2026: - Agents' Last Exam: long-horizon professional tasks, <1% average pass rate on the hardest tier (arxiv.org/abs/2606.05405) - Continual Learning Bench: do agents improve with experience? (arxiv.org/abs/2606.05661) - JudgmentBench: expert attorney preference judgments for ranking legal AI output (arxiv.org/abs/2605.25240) - SkillOrchestra: routing tasks by agent competence and cost (arxiv.org/abs/2602.19672) - SlopCodeBench: agent-written code gets more bloated and eroded with every iteration (arxiv.org/abs/2603.24755) Congrats to @amanda_dsouza, @vincentsunnchen, @chris_m_glaze, @RamyaRamakri, @fredsala, Charles Dickens, and to all our collaborators!
3
14
74
4,964
Thanks for the great conference + session @modal ! So much whitespace to cover on both public & private internal evals, as our ability to rigorously *evaluate* AI lags its advancement rate for the first time in AI history
The line to attend the evals track at Modal’s Runtime conference was practically out the door. This is why! @ajratner
2
30
2,384
Alex Ratner retweeted
Super cool work on automating benchmark creation. Check out our earlier work on this with @amanda_dsouza and folks at @SnorkelAI. arxiv.org/pdf/2510.25039v1 We showed LLMs are pretty good at tuning the difficulty knobs and creating/adapting benchmarks, as long as there is a feedback loop.
⏱️⚙️Introducing *AutoBenchmark* 📊🏁 - Creating benchmarks automatically - Benchmarking benchmark creation - Studying the role & impact of humans in the loop Blog post: facebookresearch.github.io/R… Key takeaways: 1) We find human-agent collaboration gives big wins over agents alone - fine-grained feedback in ideation stage crucial - autobench can be used to measure this in the future with stronger agents 2) We show that it's possible to make *autoresearch benchmarks* for AI research using this recipe – full recursive improvement loop! 3) Important Ingredients: Autobenchmark creation works best with feedback from two sources: benchmark solvers + external verifiers (human+AI). 🧵1/5
4
24
1,801
Excited to see @GOrlanski's new benchmark LibraryDesignBench measuring how agents can build code libraries for other agents! And excited to support via @SnorkelAI Open Benchmarks Grants
Agents will soon replace humans as the main users and designers of libraries. A good library will let future agents write correct programs with less code. Introducing LibraryDesignBench: one agent designs a library, and other agents write programs with it.
1
11
1,011
Alex Ratner retweeted
awesome discussion with @jerryjliu0 and a great group! one theme: data must be developed to be useful users/enterprises have a trove of interesting & valuable raw data (e.g. patient records, historical images/documents), but making that ready for evaluation and training requires a new form of engineering & research. this requires new interfaces & programming models involving subject matter experts to make the resulting data & environments sufficiently difficult/fair/complex!
Yesterday I hosted a fun dinner conversation with @vincentsunnchen from @SnorkelAI on evals and RL environments. The "data and RL env" companies (like Snorkel) have seen massive growth in the past few years. There's been an explosion of interest in evals. At the same time, models are ripping through benchmarks with each new release. We talked about the evals everyone is defining, what evals are still left unsolved, what’s left up to frontier models vs. intelligence that you own, and more: - A big challenge for building RL environments is “fairness” - when the model fails on a given environment, can you attribute it to the input, harness, or reward model? - Building proper rewards is hard. Some tasks are not easily quantifiable. You also want to discourage reward hacking. At the same time, you don’t want to be too prescriptive with intermediate rewards. - Long horizon evals are still extremely hard, some business processes can take up to weeks or months before the final outcome - Most regulated industries still need human in the loop to guarantee ~100% accuracy, “80%” accuracy is not good enough - As models get more intelligent, there will be a barbell of boutique data vendors (e.g. any SMB) any scaled up data providers. - Models still exhibit “jagged intelligence” where they still fail on a long tail of edge cases. - There might always be opportunities to gather unique data for a given task and posttrain models for lower cost and higher accuracy. This marks #003 in our founder dinner series. What topic should we discuss next? Let us know your thoughts below!
4
17
1,921
Strongly agree here. And a few quick thoughts: - (1) Basic research will increasingly be about the empirical science of studying and understanding, not just improving model performance. As AI models are increasingly *grown* as much as they are engineered, and as the complexity scale increasingly leads to emergent properties over discretely engineered ones, a lot of basic research is going to look more like classical empirical science - i.e. doing experiments, exploring, understanding, modeling, etc. (One type of experiment: a benchmark, if done correctly; that's what we are attempting to fund and support through @SnorkelAI Open Benchmarks Grants!) - (2) Basic research can prioritize *global optimization* (e.g. trying new architectures / algorithms / paradigms), which will always involve sacrificing local optimality at first. These first, less optimal steps - which could take months, or years - are hard for *any* commercially-funded organization to do fully (especially once basic economic gravity returns in some areas of the space). This is the gap that basic research institutions have always- and will always- need to fill! Also: these global optimization explorations do not always need 10K+ GPUs!! In fact, many will be selected specifically to explore new paradigms/approaches that are more compute efficient. - (3) Basic research has many other beneficial externalities, e.g. supporting training/teaching, open dissemination of knowledge, etc.
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
6
52
6,290
Alex Ratner retweeted
A fun end to the week as our NYC team settles into a new space 🏙️
6
35
1,774
Alex Ratner retweeted
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions. CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks. With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks. We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed. Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size. 📄 Blog: contrastive-lm.notion.site 💻 Code: github.com/Contrastive-LM/CL… 🗣️ Discord: discord.gg/5dAQEDJBs 🤗 Data & Models: huggingface.co/Contrastive-L… More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
189
638
5,540
1,054,234
Alex Ratner retweeted
can an AI generate novel philosophical ideas? there is no "unit test" for philosophy - and we need new methods to provide rigorous evaluation for AI performance in philosophy we are thrilled to support PhilosophyBench and partner with @michaelcheng76 @StanfordAILab via @SnorkelAI's Open Benchmarks Grants. if you're a philosopher and would like to participate, please reach out!
1/ Introducing PhilosophyBench from @StanfordAILab @StanfordHCI, the first independent, large-scale benchmark for evaluating AI’s philosophical capabilities. philosophybench.org
2
13
36
2,740
Alex Ratner retweeted
great benchmarks need (A) a strong thesis on which capability they measure and why it belongs in a model card and (B) the methodological and task-level rigor (including expert & programmatic QA) to back this up Terminal-Bench, Agents' Last Exam, OSWorld, and Harvey LAB are great examples of this, and we were humbled to stand behind these teams via Open Benchmarks Grants (benchmarks.snorkel.ai) - with a lot more TBA! if you're building an open benchmark - please reach out! DMs open :)
if you are building a product using AI, you should be spending >25% of your time making benchmarks and trying to get the model labs to care about said benchmarks easiest path to accelerate your progress as a company
6
37
3,482
Alex Ratner retweeted
Will be at NeurIPS '26! Congratulations to @jiayuwang111 and @ming5_alvin
Agent skills are powerful. But can they help us with agent orchestration and routing? We show they can! 🎼 Introducing SkillOrchestra - a skill-aware alternative to RL-based agent orchestration: 💰 Higher efficiency 📉 Orders-of-magnitude lower training cost 🧠 More interpretable decisions TL;DR: Stop training bigger orchestrators and start modeling skills. Paper: arxiv.org/abs/2602.19672 [1/n]
1
6
50
2,294
Alex Ratner retweeted
The full Runtime agenda is live, including talks from: @ScottWu46, Co-founder & CEO @cognition @CompleteSkeptic, Co-founder & CEO @typesafeai @dylan522p, Founder & CEO @SemiAnalysis_ @sarahookr, Co-founder & CEO @adaptionlabs @ajratner, Co-founder & CEO @SnorkelAI
9
20
177
69,022
Our team is proud to announce our investment in @SnorkelAI, the frontier AI data lab building the data and environments behind advanced AI systems. Snorkel announced a $350M Series E at a $3.5B valuation. Learn more and hear from CEO Alex on TBPN (1:46): lnkd.in/gD8aUHnb
3
8
316
Excited for this chat!!
New session 🚨 Lonne Jaffe (@insightpartners) and @ajratner (@SnorkelAI) on what's really changed in AI — and what it means to back and build companies right now. RSVP: scaleup.events/register/
2
21
1,216
Alex Ratner retweeted
Congratulations to @ajratner and the @SnorkelAI team on announcing their $350M Series E co-led by Insight Partners! Read more here: snorkel.ai/press/snorkel-ai-…
I'm excited to announce @SnorkelAI's $350M Series E at $3.5B, led by @insightpartners and @S32_VC. We've grown 18x+ in the last 12 months since launching our Data-as-a-Service offering, passing $375M ARRR this week. As AI advances to superhuman capabilities, AI data & environment development must advance with it - and basic staffing and crowdsourcing approaches are not enough. AI progress now requires deep research and technology work that combines human expertise with specialized AI in compounding ways. @SnorkelAI is building the RSI data engine and frontier data lab for this next phase. We're honored to have the support of existing investors Addition, @lightspeedvp, @GreylockVC, @GVteam, P7, Factory, @WellsFargo, Walden Catalyst Ventures, and new investors @ThirdPointLLC, @MarchCPs, @BlumbergCapital, @AllegisCapital, @Frontlinevc, and @standard_vc. – @SnorkelAI started as a research project a decade ago at @StanfordAILab. Our thesis was simple: AI progress would become increasingly data-centric – and therefore data development should be studied as a true research and technology problem, not just a staffing and crowdsourcing one. Today, as AI capabilities verge on superhuman, building the data and environments to safely measure and train AI is becoming too hard for even the smartest human experts to do alone. Only humans and AI agents, collaborating together in compounding ways, can meet the accelerating needs of the frontier, and keep humans in the driver’s seat of AI progress for decades to come. At @SnorkelAI, we are building the data lab to define the shape of this new “Data 2.0” frontier, and the new paradigms of human-computer interaction needed to advance it. Our key focus is building the RSI engine for data, where specialized AI models accelerate and improve human expert output, and in turn, scaled human supervision is used to continuously evaluate and improve these models – creating a powerful compounding loop to keep pace with an accelerating RSI frontier. With this round of funding, we are also doubling down on our commitments to support data development for open benchmarking and evaluation (more news here soon!); an increasingly diverse ecosystem of general and specialized intelligence; and a path to safe, well-aligned AI built on robust training and evaluation data. Data development will guide and drive the next stages of AI – and must do so in a human-centric, AI accelerated, open, diverse, and safe way. We are excited to support this mission in the next decade of research ahead at @SnorkelAI. More thoughts here: snorkel.ai/blog/data-2-0-and…
2
15
1,910
I'm excited to announce @SnorkelAI's $350M Series E at $3.5B, led by @insightpartners and @S32_VC. We've grown 18x+ in the last 12 months since launching our Data-as-a-Service offering, passing $375M ARRR this week. As AI advances to superhuman capabilities, AI data & environment development must advance with it - and basic staffing and crowdsourcing approaches are not enough. AI progress now requires deep research and technology work that combines human expertise with specialized AI in compounding ways. @SnorkelAI is building the RSI data engine and frontier data lab for this next phase. We're honored to have the support of existing investors Addition, @lightspeedvp, @GreylockVC, @GVteam, P7, Factory, @WellsFargo, Walden Catalyst Ventures, and new investors @ThirdPointLLC, @MarchCPs, @BlumbergCapital, @AllegisCapital, @Frontlinevc, and @standard_vc. – @SnorkelAI started as a research project a decade ago at @StanfordAILab. Our thesis was simple: AI progress would become increasingly data-centric – and therefore data development should be studied as a true research and technology problem, not just a staffing and crowdsourcing one. Today, as AI capabilities verge on superhuman, building the data and environments to safely measure and train AI is becoming too hard for even the smartest human experts to do alone. Only humans and AI agents, collaborating together in compounding ways, can meet the accelerating needs of the frontier, and keep humans in the driver’s seat of AI progress for decades to come. At @SnorkelAI, we are building the data lab to define the shape of this new “Data 2.0” frontier, and the new paradigms of human-computer interaction needed to advance it. Our key focus is building the RSI engine for data, where specialized AI models accelerate and improve human expert output, and in turn, scaled human supervision is used to continuously evaluate and improve these models – creating a powerful compounding loop to keep pace with an accelerating RSI frontier. With this round of funding, we are also doubling down on our commitments to support data development for open benchmarking and evaluation (more news here soon!); an increasingly diverse ecosystem of general and specialized intelligence; and a path to safe, well-aligned AI built on robust training and evaluation data. Data development will guide and drive the next stages of AI – and must do so in a human-centric, AI accelerated, open, diverse, and safe way. We are excited to support this mission in the next decade of research ahead at @SnorkelAI. More thoughts here: snorkel.ai/blog/data-2-0-and…
73
115
412
89,311
Congrats on yet another great milestone, @SnorkelAI! 7.5 years ago and long before the recent AI wave, we @GVteam were honored to co-lead Snorkel's very first round of funding alongside @GreylockVC (@saammotamedi + @saranormous), backing a stellar team of founders coming out of Stanford's AI lab. It has been incredible to see what @ajratner and the entire team have accomplished in that time, including today’s $350M round at a $3.5B valuation, while surpassing $375M in ARR.
I'm excited to announce @SnorkelAI's $350M Series E at $3.5B, led by @insightpartners and @S32_VC. We've grown 18x+ in the last 12 months since launching our Data-as-a-Service offering, passing $375M ARRR this week. As AI advances to superhuman capabilities, AI data & environment development must advance with it - and basic staffing and crowdsourcing approaches are not enough. AI progress now requires deep research and technology work that combines human expertise with specialized AI in compounding ways. @SnorkelAI is building the RSI data engine and frontier data lab for this next phase. We're honored to have the support of existing investors Addition, @lightspeedvp, @GreylockVC, @GVteam, P7, Factory, @WellsFargo, Walden Catalyst Ventures, and new investors @ThirdPointLLC, @MarchCPs, @BlumbergCapital, @AllegisCapital, @Frontlinevc, and @standard_vc. – @SnorkelAI started as a research project a decade ago at @StanfordAILab. Our thesis was simple: AI progress would become increasingly data-centric – and therefore data development should be studied as a true research and technology problem, not just a staffing and crowdsourcing one. Today, as AI capabilities verge on superhuman, building the data and environments to safely measure and train AI is becoming too hard for even the smartest human experts to do alone. Only humans and AI agents, collaborating together in compounding ways, can meet the accelerating needs of the frontier, and keep humans in the driver’s seat of AI progress for decades to come. At @SnorkelAI, we are building the data lab to define the shape of this new “Data 2.0” frontier, and the new paradigms of human-computer interaction needed to advance it. Our key focus is building the RSI engine for data, where specialized AI models accelerate and improve human expert output, and in turn, scaled human supervision is used to continuously evaluate and improve these models – creating a powerful compounding loop to keep pace with an accelerating RSI frontier. With this round of funding, we are also doubling down on our commitments to support data development for open benchmarking and evaluation (more news here soon!); an increasingly diverse ecosystem of general and specialized intelligence; and a path to safe, well-aligned AI built on robust training and evaluation data. Data development will guide and drive the next stages of AI – and must do so in a human-centric, AI accelerated, open, diverse, and safe way. We are excited to support this mission in the next decade of research ahead at @SnorkelAI. More thoughts here: snorkel.ai/blog/data-2-0-and…
2
5
37
4,276
Alex Ratner retweeted
“As AI systems take on longer-running, higher-stakes work, the data needs from frontier AI labs and enterprises are shifting from generalist data that humans can write down to specialized data, environments, and evaluations that often need to be built.” In the below post, @lonnej and I share some thoughts about our (@insightpartners's) recent investment in @SnorkelAI and what got us so excited to partner with @ajratner and the Snorkel AI team as they build the data lab for “Data 2.0”. insightpartners.com/ideas/be…
4
14
1,089