Building ai systems, Stanford cs phd @hazyresearch, Incoming assistant professor @caltech, Leading the frontier performance research team @togethercompute

California
AI has been built on one vendor’s stack for too long. AMD’s GPUs now offer state-of-the-art peak compute and memory bandwidth — but the lack of mature software / the “CUDA moat” keeps that power locked away. Time to break it and ride into our multi-silicon future. 🌊 It's been a blast working with the amazing @_williamhu, @Drewwad and team; we present HipKittens!
14
95
582
229,998
Simran Arora retweeted
Inspired by this, I have spent a lot of time creating the same version but for Apple's GPU. - part 1: basically @Si_Boehm 's worklog but Metal lenguyen.vercel.app/note/met… - part 2: simdgroup matrix and TensorOps (Metal 4 API) lenguyen.vercel.app/note/met…
we got early access and hipkittens is now running on amd helios mi455! here we share an educational gemm ladder to demonstrate how to use the new helios features. overall, we found that the patterns identified in hipkittens for writing performant mi350/355 kernels translate well to helios making the forward port quite smooth!
1
2
5
367
Simran Arora retweeted
This thing has been incredible for us @togethercompute and we are releasing in 48hrs for all of you.
TogetherLink, coming soon.
3
7
56
10,386
🇺🇸✈️💫 llmh is at the frontier of deploying ai across the economy! super exciting!!
Today, Long Lake completed our $6.3B acquisition of Amex GBT. Long Lake acquires and transforms generational businesses with AI across the American services economy. Amex GBT is the travel partner for 17,500 businesses in 140 countries. Last year, Amex GBT booked 35 million trips for 10 million travelers. This marks Long Lake’s 40th acquisition, and our family of companies now employs nearly 30,000 people. We founded Long Lake three years ago with the thesis that AI is going to change every company, but there’s a large overhang between AI capabilities and how most businesses leverage AI tools. We partner with strong, profitable, and growing businesses to help them deploy AI and improve their customer service. Our first cohort of companies has doubled their EBITDA in less than two years through topline and productivity growth, all while increasing headcount. (We’re proud to say that we’ve never conducted a layoff.) We’re excited to share a bit more about what we’ve been up to.
39
3,286
💫💫
Open collaboration is moving GPU kernel development forward. AMD and contributors from @Caltech, @Stanford, @togethercompute and the broader developer community teamed up to bring HipKittens support to AMD Instinct MI455X GPUs. The collaboration produced peak-performance GEMM kernels for the Helios GPU architecture using HipKittens. The walkthrough starts with a baseline kernel and progresses through optimizations including asynchronous memory transfers, Tensor Data Movers (TDM) and workgroup-cluster multicast. Developers can follow the code and see how each technique shapes kernel performance. Explore the technical breakdown and code: rocm.blogs.amd.com/software-…
4
36
5,669
Simran Arora retweeted
Open collaboration is moving GPU kernel development forward. AMD and contributors from @Caltech, @Stanford, @togethercompute and the broader developer community teamed up to bring HipKittens support to AMD Instinct MI455X GPUs. The collaboration produced peak-performance GEMM kernels for the Helios GPU architecture using HipKittens. The walkthrough starts with a baseline kernel and progresses through optimizations including asynchronous memory transfers, Tensor Data Movers (TDM) and workgroup-cluster multicast. Developers can follow the code and see how each technique shapes kernel performance. Explore the technical breakdown and code: rocm.blogs.amd.com/software-…
4
24
107
13,146
Simran Arora retweeted
my journey into kernel engineering began with porting @Si_Boehm’s gemm ladder to amd - definitely feels like a full circle moment!
Replying to @simran_s_arora
our blog is inspired by the OG @si_boehm’s educational gemm ladder! check out our blog here: rocm.blogs.amd.com/software-… with @neoblizzz, ryan swann, @_williamhu, @seansiddens, @drewwad, julia zhang, alex underwood, alex dutu, @dylan__lim. thank you @amd @AIatamd for the support!
3
58
7,636
we got early access and hipkittens is now running on amd helios mi455! here we share an educational gemm ladder to demonstrate how to use the new helios features. overall, we found that the patterns identified in hipkittens for writing performant mi350/355 kernels translate well to helios making the forward port quite smooth!
Made with AI
2
24
157
14,395
key new helios features include asynchronous HBM to shared memory data movement using DMA (TDM), a simplified cache structure (fewer cache levels) and thread hierarchy (32 threads per wave), fine-grained synchronization mechanisms, and workgroup-cluster launch and multicast.
1
2
28
1,989
💫😸💫
ThunderKittens is now on Vera Rubin! It's been a fun past week playing with new kernels and hardware. Check out the blog below for more details on how we pushed our GEMMs to cuBLAS levels of performance! 🐱⚡️🔭
1
25
2,314
Simran Arora retweeted
ThunderKittens now runs on Rubin GPUs!
thunderkittens is now running on nvidia vera rubin! our kernels team got early access to nvl72 and spent the past few days digging through the new isa and bringing nvfp4 + fp8 gemms to life after reworking the kernels for rubin, we pushed them past 22 and 12 pflops respectively — competitive with cublas + cute dsl
1
4
55
4,306
thunderkittens is now running on nvidia vera rubin! was a smooth porting experience with some fun new updates -- checkout the blog! great work by @dylan__lim, Xinyi and the @togethercompute team!
thunderkittens is now running on nvidia vera rubin! our kernels team got early access to nvl72 and spent the past few days digging through the new isa and bringing nvfp4 + fp8 gemms to life after reworking the kernels for rubin, we pushed them past 22 and 12 pflops respectively — competitive with cublas + cute dsl
3
18
121
9,896
Simran Arora retweeted
We are partnering with @Equinix and @nvidia to build a global inference fabric for inference. This partnership brings inference (literally) close to enterprise data and applications through our new inference edges colocated in Equinix's distributed, enterprise-class, low-latency datacenters.
Open source is becoming the default way enterprises build with AI. Not just which model runs, but where it runs. Today we're partnering with @Equinix and @nvidia on Equinix Inference Exchange: our open-model platform, live across Equinix's global data centers. Model choice and performance were never a trade-off. newsroom.equinix.com/2026-09…
8
14
112
19,689
Simran Arora retweeted
the hipkittens have gone distributed... meet Moey 😼 more soon!
Made with AI
10
2
34
1,563
Simran Arora retweeted
We executed a landmark partnership with @HUMAIN today to offer open source models on 250 mw (-100K+ chips) of capacity over the next 12 months. The demand for open and custom AI tokens is vertical and this partnership represents one of the largest so far focused on delivery OSS tokens globally.
We just signed one of the largest AI infrastructure deals for open source, period. 250MW data center. $5B+ in annualized revenue. Built with @HUMAIN in Saudi Arabia. tinyurl.com/4nzcjcu7
9
13
137
30,157
Cool paper! Multiple passes over the sequence with fixed state models is a powerful axis!
In-context continual learning requires models to accumulate experience and reuse it later in the same sequence. But an RNN compresses an ever-growing history into a fixed-size state, where each token gets a single write into memory. We study dynamic compression: letting the model revisit the past and reorganize its state as it discovers what needs to be reused.
3
3
73
8,782
Simran Arora retweeted
Check out Hawkeye, which writes high-performance kernels utilizing advanced architectural features (e.g., TMA for async data transfer on Blackwell, L2 locality on MI350) from only one handwritten example and ~10 unit tests. Hawkeye can port kernels across architectures (Ampere, Hopper, Blackwell), chips (NVIDIA, AMD), and precisions (FP8, NVFP4, MXFP4). Great work, co-led by @AryaTschand and @keramakr!
We’ve seen an explosion of new ML chips with unique architectural features, but software support remains the critical bottleneck Achieving peak performance increasingly relies on hardware-specific optimizations in the kernels, but we observe that coding agents are particularly weak at this Introducing Hawkeye, a framework that brings hardware-awareness to coding agents by grounding them in a minimal and comprehensive taxonomy of optimization strategies For new GPU or ML accelerator architectures, you only need to write 10 unit tests and solution kernels (one per optimization strategy), and we show that coding agents can effectively scale test-time compute with this minimal supervision to write hardware-aware kernels Hawkeye can port kernels across architectures (Ampere, Hopper, Blackwell), vendors (NVIDIA, AMD), and precisions (FP8, NVFP4, MXFP4) while consistently leveraging hardware features and approaching expert kernel performance Work co-led with @keramakr and done in collaboration with Alexander Ingare @simonguozirui @18jeffreyma @ZishenW @simran_s_arora @Azaliamirh @profvjreddi
2
17
209
18,064