Asst. Prof @PrincetonCS, Chief Scientist @togethercompute. Machine learning & systems.

Stanford, CA
FlashAttention is widely used to accelerate Transformers, already making attention 4-8x faster, but has yet to take advantage of modern GPUs. We’re releasing FlashAttention-3: 1.5-2x faster on FP16, up to 740 TFLOPS on H100 (75% util), and FP8 gets close to 1.2 PFLOPS! 1/
35
344
2,249
355,541
Tri Dao retweeted
We're releasing tev1-4B-experimental, a Jev-like classifier finetuned on top of Qwen3.5 4B. We're making it available on Together serverless at $0.042/M input & $0/M output. Also releasing the data recipe & a tutorial on how to finetune your own (Tev1 cost $17 to train!).
43
82
1,284
136,271
Tri Dao retweeted
we got early access and hipkittens is now running on amd helios mi455! here we share an educational gemm ladder to demonstrate how to use the new helios features. overall, we found that the patterns identified in hipkittens for writing performant mi350/355 kernels translate well to helios making the forward port quite smooth!
Made with AI
2
24
157
14,357
Mayank put in a crazy amount of work to get pretraining to work on 3 gens of Nvidia GPUs and 2 gens of TPUs! Very good model for such a small size
We pretrained a 2.3B MoE (360M active) Hybrid Mamba-2 that lands within a few points of Llama-3.2-3B using <1% of its pretraining FLOPs. No dedicated cluster. The run hopped between H100s, A100s, V100s (yes, V100s) and TPU v5p/v6e on a single codebase. Meet Rigel 🧵
7
17
282
26,208
A lot more gpus coming for open models
We just signed one of the largest AI infrastructure deals for open source, period. 250MW data center. $5B+ in annualized revenue. Built with @HUMAIN in Saudi Arabia. tinyurl.com/4nzcjcu7
12
6
211
28,842
These guys move fast, 1st rack already ships. More inference compute is always welcome
We've raised $700M at a $21B valuation from Jane Street, Kleiner Perkins, Sequoia, A16Z, Peter Thiel, BCV, and Blackstone. We're also excited to share that we've shipped our first rack to Jane Street.
9
4
180
38,441
I really like this new benchmark. Has the flavor of ARC-AGI3 but it’s pure text so you don’t have to worry about the vision capability
We’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our novel discovery games and seeing how they score. Each game is a text-based environment, so they probe discovery capabilities in the natural domain of language models, rather than requiring additional, potentially confounding, visual understanding. TL;DR frontier models have improved a lot over the last few months. But they are still stumped by some surprisingly simple problems, even in their native text domain. With @cocosci_lab (@Princeton) @MITCoCoSci (@MIT) @SchmidhuberAI (@KAUST_News) @misovalko (@Inria) @tri_dao (@PrincetonCS) @RMBattleday @zebkDotCom @FraserGreenlee @akaijsa @ClareMaguire @TimMuller1 @kubicek_ales @physicscat0x7d @SukritSumant @thoughtchannel_ (1/5)
1
16
211
33,408
Tri Dao retweeted
the gains come entirely from wrapping π0.5 in an agent that plans, verifies, and recovers. no new finetuning. @lianegalanti and team make the case that "agents for robots" can narrow capability gaps in off-the-shelf robot policies.
Robot policies can move but can't think. LLMs can think but can't move. So we connected them. Real robot: 16.7% → 97.3% Sim (LIBERO-PRO): 12.8% → 53.3%
3
3
27
5,161
Congrats to this super ambitious team
Today we’re announcing our Series B. We’ve raised $200M at a $2B valuation from Greenoaks with participation from Index Ventures, Hanabi, A*, Bain Capital Ventures, CVS Health Ventures, and Definition. Our mission is to simulate all eight billion people on earth, accurately.
5
6
84
18,994
Putting LLM brain on robots -> 4x SOTA with no extra training. I’ve been very surprised by how well this works. The time for agents running on robots is coming soon
Robot policies can move but can't think. LLMs can think but can't move. So we connected them. Real robot: 16.7% → 97.3% Sim (LIBERO-PRO): 12.8% → 53.3%
12
27
272
48,835
Tri Dao retweeted
Kimi K3 is now live on Together AI. We’re proud to be a Day 0 launch partner for @Kimi_Moonshot’s open frontier model, built for long-running agentic workflows across code, tools, vision, and research.
21
21
160
175,062
Congrats to this team, really strong on both hardware design and kernels
We’ve raised $300M in Series C funding at a $10.3B valuation from Sequoia, Andreessen Horowitz, Jane Street, Argo, and SK Hynix. Our mission is to run the world's inference. This round accelerates production of our inference clusters. We've opened an 80,000-sqft, 10-MW facility 15 minutes from our office to expedite production and prototyping.
2
7
215
27,016
Tri Dao retweeted
We're introducing Provisioned Throughput: reserved inference capacity for frontier open models, with token-based pricing and a 99% uptime SLA. Serverless simplicity, guaranteed capacity, up to 90% lower cost vs. Opus 4.8. Get started with MiniMax M3 + GLM-5.2, read more 🧵
11
8
92
27,762
We’re serving 400T tokens / month and the demand for open models just keep going up
We @togethercompute believe intelligence should be abundant, not expensive. Today we announced our Series C funding of $800m @ $8.3B valuation, to continue to build the world's most efficient platform for generative AI. Thanks @nikogallogly for telling our story in @nytimes! shorturl.at/SooOP
5
11
191
26,088
Tri Dao retweeted
We @togethercompute believe intelligence should be abundant, not expensive. Today we announced our Series C funding of $800m @ $8.3B valuation, to continue to build the world's most efficient platform for generative AI. Thanks @nikogallogly for telling our story in @nytimes! shorturl.at/SooOP
67
87
488
175,166
It's wild how quickly Etched designed and got the chips out, all within 2 years. They went deep, hardcoding attention into silicon and getting very high MFU. This kind of hardware tailored made for LLM inference is soon gonna bring cost of intelligence down 10x
We're coming out of stealth. We've built our first racks after a successful A0 tapeout, $1B+ in customer contracts, and $800m raised. Early customer tests show us achieving SOTA throughput, latency, and power efficiency on inference workloads. Our first racks ship this summer.
26
59
1,151
154,347
Tri Dao retweeted
Samir Menon @blintzbase and I are thrilled to announce Sail @sailresearchco ! We build infrastructure for long-horizon agents: inference served at unbeatable prices-per-token for open models, plus sandboxes designed to run for days, weeks, or longer. We've raised $80M, w/ our seed led by @Sequoia and series A led by @KleinerPerkins. We're using this capital to build the most efficient infrastructure for long-horizon agents. What makes agents so different? Unlike a human waiting at a keyboard (top priority: speed), agents need scale, reliability, and sustainable cost. Sail finds this efficiency everywhere in the stack: we carefully choose our chips, write custom inference engines, and run a global controller that fully utilizes every computer in our fleet. Tight integration from silicon to API lets Sail open up the cost / latency frontier to our customers - the most patient agents can now access 10x more intelligence per dollar. We're excited to be working with great companies like @parallelweb, @detaildotdev,@Jackandjillai, and @quadrillion_ai to deploy long-horizon agents with trillions of tokens. Our team is thoughtful in our engineering craft and relentlessly ambitious in our pursuit of peak performance. We previously trained at companies like NVIDIA, OpenAI, Google, and so many trading firms. Now we're ready to do the work that will define our careers, in the most compute intensive market of all time. Welcome to the era of abundant intelligence. We can't wait to build with you!
104
63
548
262,474
Tri Dao retweeted
Together with my co-founders Michael @MichaelPoli6, Stefano @Massastrello and Armin @athmsx, I am excited to announce @RadicalNumerics is emerging from stealth with a $50M seed round to build general biological intelligence. We’re also sharing an early preview of our new model Omnii, the most powerful genome language model to date. Omnii preview link: radicalnumerics.ai/blog/radi… At Radical Numerics, our mission is to master the code of life, and to drive the frontier of biological AI for both design and defense. This is our dual mandate, which comes from something our own team helped make possible. Our founding team trained Evo and Evo 2, the largest biological AI models (40B params) trained on DNA sequences. Trillions of tokens across all of life, from microbes to mammals. It’s fully open source, and created the field now known as generative genomics. Last year, scientists used Evo to generate the world’s first complete genome from scratch using AI. Turns out it was a bacteriophage—a type of virus. It functioned in the real world, and in this case it was harmless. But for us, it was a clear turning point. It showed that AI is no longer just analyzing biology. It is on the cusp of generating functional lifeforms. Eventually, AI will have the power to design and control life itself. That should make all of us incredibly excited, and incredibly uneasy. (Anyone can design DNA with a new function, and have it synthesized and delivered, like something from Amazon Prime). The same technology that will help us cure cancer is the very technology that might create the next global pandemic, or worse, allow the creation of bioweapons that can wipe out populations. We believe these forces are inseparable. If you work on the frontier of biology, you have to build technology to safeguard it from its misuse. Existing biosecurity tools are sorely losing the arms race, relying on outdated “have I seen this exact thing before?” style algorithms. We founded Radical Numerics to turn the tide. And we can’t do that by training on textbooks and natural language. We must understand the language of biology from the raw physical data itself, to reason across every molecule and modality, from DNA to proteins. The next frontier for AI goes far beyond chatbots or video generators to models that can understand and engineer life. Today, we’re previewing Omnii, which is already far surpassing Evo 2, and will continue improving as we scale and add new modalities (training now). 1. For human health, Omnii can read and write whole genomes (more on writing later). It’s state of the art (SOTA) on detecting causal variants for disease, and can rank Alzheimer's mutations zero-shot. We’re partnering with a diagnostics company to use Omnii for early cancer detection (pancreatic and multi-cancer). 2. For defense, Omnii is SOTA at detecting AI-generated pathogens. We benchmarked existing detection tools, and they simply can’t detect the AI-generated ones (“deepfake viruses”). We’re partnering with a US national lab to pilot Omnii for detecting the next pandemic, both natural and AI-generated. We have a data center full of Blackwells in construction now to build the most powerful biological AI models ever. This mission takes a new kind of AI lab that can actually scale on physical, biological data: new alignment research (mid/post training), scaling long context, building out mech interp teams to dissect what these models learn, new architectures and systems designs, all from the ground up. Our team is made up of AI researchers and scientists from top labs and institutions (e.g. Stanford, MIT, Google DeepMind), but more importantly, we all share the belief that this is the most important challenge of our lifetime. If you feel similarly, we are hiring. We aim to bring the brightest minds in AI and science together to save lives. Thanks to our partners on this journey, led by Emergence Capital @emergencecap, with Obvious Ventures @obviousvc, Triatomic @TriatomicCap , and Patrick Collison @patrickc. Our advisors include Eric Horvitz @erichorvitz, CSO of Microsoft, Chris Re @HazyResearch of Stanford, George Church @geochurch of Harvard, and Andrew Weber @AndyWeberNCB, former Assistant Secretary of Defense for Nuclear, Chemical and Biological Defense Programs. Fortune article: fortune.com/2026/06/15/exclu… Jobs: radicalnumerics.ai/join-us
154
297
1,383
2,577,586
Tri Dao retweeted
Cartesia Sonic 3.5 is now available on Together AI. We added 150+ @cartesia Sonic 3.5 voices to voice finder, so developers can listen, compare, and pick the right voice for real-time agents before deploying on Together AI.
4
4
29
5,846