Systems and GPU Performance Mechanic - TBD Ex. CUTLASS 3.x / 4.x etc

As of last week, I am no longer at NVIDIA 🧵 Leaving the CUTLASS team was extremely hard. I will dearly miss my incredible colleagues and the extremely compelling mission statement of creating the world's best accelerator programming model w/ hardware software codesign 💚
16
17
370
28,680
Vijay retweeted
i shed a tear reading this blog post by a deepseek kernel engineer
64
463
6,747
210,078
Vijay retweeted
Last month I wrote about how we can build a positive and safe future for everyone: meta.com/thefutureisforevery… Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. The reality is: - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind. - Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well. Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built. - Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators. - Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well. I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
1,928
2,861
28,449
7,751,607
Vijay retweeted
Introducing Muse, the personal agent that understands your goals and works 24/7 to get things done for you.
3,481
2,554
36,743
8,679,507
1/ today we’re releasing muse spark 1.3—available in muse code & the meta model api. this is our most capable model yet—frontier performance almost too cheap to meter. much stronger at agentic and coding with better usability. we think users will really notice the jump.
200
306
3,608
654,055
Vijay retweeted
it’s easy to trick yourself into thinking you’re a net producer of content by posting. but unless you’re posting ai generated blogs on how you (not ai ofc) wrote a kernel to beat cublas and cudnn for one very important shape with definitely not undertuned baselines, you’re not creating
it's easy to trick yourself into thinking you're a net producer of content by tweeting. but unless you're producing long-form content, you're not creating
5
1
84
6,142
Vijay retweeted
Today we’re officially launching @explabsai on YC. Companies spend more on AI every month, but none of it becomes an asset they own. Up to 97% cheaper and 50% higher quality.
150
51
369
43,433
Vijay retweeted
Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.
1,423
1,279
14,845
3,230,474
Vijay retweeted
we just open sourced a tool train a specialized model at Fable quality and half the cost it uses agent traces to continuously improve via • distillation and RL • model routing • token compaction
21
21
188
25,824
Vijay retweeted
(1) Today we're releasing Muse Spark 1.1 -- a strong agentic and coding model at a very low price. It's available through our new Meta Model API and in Meta AI.
5,355
3,582
45,466
23,685,299
Vijay retweeted
After some mathematical rewrite, turns out all of transformer is a series of gemm + epilogue. Given a few optimized primitives, LLMs (and novice humans) can write speed-of-light kernels for all transformer ops!
LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes them to hide in the matmul’s shadow, fused into its epilogue before results leave the chip. Bonus: LLMs can write fast CODA kernels too (approaching SoLs).
17
128
1,207
135,502
Vijay retweeted
LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes them to hide in the matmul’s shadow, fused into its epilogue before results leave the chip. Bonus: LLMs can write fast CODA kernels too (approaching SoLs).
16
102
688
205,388
Vijay retweeted
We built a kernel abstraction to rewrite the entire transformer stack as GEMM + Epilogue kernels! Neural net architectures such as transformers consist entirely of matrix multiplications and elementwise nonlinearities such as RMSNorm, log sum exp, and gated activations. Fusing these elementwise nonlinearities into GEMMs in both the forward and backward passes allows us to make training and prefill as compute-bound as possible! Our kernel abstraction CODA is implemented in CuTeDSL, and by abstracting away the fixed prologue and main loop of the GEMM kernel, we expose an epilogue function where LLMs like Claude can easily implement elementwise nonlinearities in fusions approaching speed-of-light!
LLM training is built on fast MatMuls. But many surrounding ops still run as memory-bound kernels. CODA reparameterizes them to hide in the matmul’s shadow, fused into its epilogue before results leave the chip. Bonus: LLMs can write fast CODA kernels too (approaching SoLs).
1
22
179
20,533
Vijay retweeted
We’ve developed our own inference engine Runtime-Optimized Serving Engine (ROSE) to serve models ranging from embeddings to trillion-parameter LLMs. With CuTeDSL integrated into our inference engine, Perplexity can build the specialized GPU kernels faster to bring models up to peak performance on NVIDIA Hopper and Blackwell GPUs.
72
112
1,039
163,454
Vijay retweeted
new walk of shame: agent still working, but the cafe closed
259
184
5,391
606,460
[ENG SUB] how it feels to use eager pytorch in 2025
28
59
472
88,876
Tomorrow: Blackwell Programming lecture by yours truly at Stanford CME213, Gates B3, 1:30–2:50 PM. Bring sharp questions.
7
9
135
7,668
Vijay retweeted
We're open-sourcing FlashKDA — our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. Achieves 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention. Explore on github: github.com/MoonshotAI/FlashK…
45
186
1,798
221,668
Vijay retweeted
Meta released Avocado, they call it Muse Spark. It's not open source (a bit sad). Meta TBD lab rebuilt the entire pretraining stack in 9 months and reached similar capability with >10x less compute than Llama 4 Maverick. I still think infra is the real moat in AI labs. You can train models much faster with a good infra, and it allows researchers to experiment with many more ideas much more quickly.
40
28
639
54,354
Excited to share what we’ve been building at Meta Superintelligence Labs! We just released Muse Spark, our first AI model. It's a natively multimodal reasoning model and the first step on our path to personal superintelligence. We've overhauled our entire stack to support scaling, and this is just the beginning. ai.meta.com/blog/introducing…
71
173
1,656
243,780