Open Source Model Lover @ NVIDIA AI Views my own.

Toronto, Ontario
Say what you want, these two pictures ARE different
since we're doing the AI conscious debate again, i want to remind everyone that despite how nice the insanely surface-level analogies seem, we would do well to scrutinize things at an appropriate level of detail this --> is not the same as this -->
1
12
1,116
idk putting anything in something called a torture pit is pretty fucked imo
6
428
Btw if you're not watching this this weekend then what are you even doing with your life 100 outta 10 event
RSI Training goes LIVE at 5 PM PT today! We’re giving each agent access to a production Slurm cluster, 1,000 GPU-hours, and 6 days to improve Nemotron 3.5 model performance. Watch it all happen live -> rsiarena.org
9
1,215
United leaving the overhead lights on for the red eye is a crazy move honestly
2
1
12
725
I don't have llm psychosis, because I've internalized that it's normal to just talk to AI like a guy
2
11
600
ITS TIME TO BUILD A PET FOR MY NET NAVI
🚨 MUSE ECOSYSTEM ALERT 🚨 today we are announcing Muse Gadgets! this is an open-source ESP32 firmware and Linux SDK for anyone to make hardware that works with Muse we are also releasing our own gadget—Muse Home Link—to enable your muse to work with your smart devices (TV, speakers, etc.)
7
445
Chris 🇨🇦 retweeted
Introducing Prime Inference: We've served trillions of tokens for RL and dedicated customer deployments To own your intelligence, you need to own your inference Unpacking our inference stack
73
101
1,128
161,561
pewds pliney crossover? Is this the prime timeline?
4
13
882
This lab is going to have trilliums of dollars of impact.
Today we're unveiling Trillium Labs @trillium_labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next. We're built around the theory of change that you need more eyes to solve hard technical problems. We have faith in the scientific methods and communities that humanity has built, and worry that AI is becoming too closed to utilize them. Trilliums are wildflowers that bloom briefly in the spring, before the forest canopies fill out. Though they are small, they lay the foundation for the cycles of growth and nourishment through the rest of the year. At Trillium Labs, the recipes will be the slow nutrients for the seasons and the model releases will be the blooms. Building an institution dedicated to this is needed because, much as nature’s trilliums are slow to expand and grow, the open-ecosystem needs time and dedicated resources to catch up. I co-founded with with a long-time friend and collaborator Tom Zick (@thesezickbeats). We're hiring (full time + student collabs/interns), we're fundraising, and we're looking for compute. Please get in touch if you're interested in helping out. Offices based in the Bay Area and Cambridge MA, remote okay. I’m in the Bay Area until for The Curve and COLM to connect with people who are interested. We’re thankful to have initial support from Halcyon Futures and Schmidt Sciences with more funding en route to enable our ambitions of scaling. Our advisors @Thom_Wolf, @HannaHajishirzi, @gneubig and @ctnzr have been instrumental to building the ecosystem that exists today, and I’m stoked to get to keep working with them.
1
17
1,197
i feel hopeful for the future
2
11
631
Chris 🇨🇦 retweeted
YouTube plans to introduce AB testing for videos. I agree with many others that this is a bad idea. Maybe you squeeze out some more watch time in the short term, but undermining a sense of a shared viewing experience among the audience is too steep a price.
102
299
11,280
361,751
No big deal, just simulating the entire world.
I joined Daniel and Chris on Practical AI to discuss open models, NVIDIA Cosmos, and how world models can help AI understand and simulate the physical world. We explore why openness matters and what it takes to build increasingly capable physical AI systems. Watch: piped.video/watch?v=W1DSS-Sz… Listen: practicalai.show/374
3
479
Big yaps about my favourite model in the world: Nemotron! We get into the what, the how, and the why now!
alright folks if you are interesting in learning more about the inner working of nvidia’s frontier-class open model with a cool name I have this 1h30 discussion with the man @llm_wizard in it we discuss: - latent moe - speed - aggressive GQA - speed - bonker linearization
4
2
23
1,040
Chris 🇨🇦 retweeted
So I would like to share my analysis of PD disagg that I've run by Horace and many folks at Thinky You may have heard that PD disagg helps with per-token "tail latencies" in ITL/TPOT, as prefill interrupts decode. This never was a convincing explanation to me though because people don't actually see single-token tail latencies during decode (what matters is how long the entire response takes - 100-10k tokens!). Insight: steady-state response tells us 1. PD disagg actually improves mean per-token latency 2. PD disaggregation helps most for prefill-heavy workloads, it doesn't help for ultra low-latency [Image 1] Here's an intuitive model. Imagine having 4 regular samplers, getting traffic so they spend 75% of their time on prefill and 25% on decode. The batch size (running requests) of each sampler is determined by a steady-state feedback loop. As new requests come in, they increase batch size, which then decreases interactivity as requests finish slower. If a decode batch takes 10ms, average TPOT would be 40ms in this case (only 25% time spent on decode). [Image 2] Now let’s say that we split it into 3P/1D disaggregated samplers. Prefill takes the same amount of time, and 1 decode worker handles 4x the number of requests. But it runs with the same batch size (throughput is equal), as it processes 4x the number of decode requests, each 4x faster. Each decode batch would take 10ms, but average TPOT has been lowered from 40ms → 10ms. In other words, the steady-state model of PD disagg is: if your regular engine is spending a substantial amount of time on prefill, then at scale, you can get "increased interactivity” with multiple 1/(Decode Time %), while keeping the exact same batch size and throughput. With 1:1 PD, you get 1/0.5 = 2x interactivity. With 1:4 PD, you get 1/0.8 = 1.25x interactivity. I've validated this in production on a bunch of different inference workloads. Basically - decode-bound (low-latency) inference does not need disagg. But if you see your inference engine is spending a good % of its time running prefill passes, it's time to bust out the PD disagg. --- Assumptions (if you want to check my work): - This is independent of any prefill delayer settings; it actually assumes "perfect" prefill/decode scheduling on the regular engines. - This assumes that the network fabric for sending P->D kv caches is not your bottleneck - e.g., 100 GB/s Infiniband can transfer KV caches in 100 ms. - Obviously if you do PD disagg, you should also specialize engine configurations - like prefill may use EP4 while decode uses wider EP32 topologies. This gives you additional speedup! But this is secondary. - I don't include mixed chunk in this analysis because mixed-chunk has not very good support, e.g., it's incompatible with speculative decoding (very important optimization) in most engines. - P/D disagg also affects the amount of usable GPU HBM you get for storing persistent KV prefix caches -- in this analysis you lose total KV space if prefill takes <50% of your time. I think this is largely mitigated by tiered caches + Mooncake in practice.
I find that surprisingly few people have *really* thought about why PD disagg is a good idea. Even fewer people have *really* thought about how PD disagg compares to mixed chunk prefill (as Matthew describes here).
24
51
483
67,478
I need people to wake up to how bad benchmarks are faster - how can I do this? What do I need to do? Paper? Blog? Interview Xeo?
12
26
1,664
Chris 🇨🇦 retweeted
We just open-sourced the world's fastest WebGPU kernels for 200+ machine learning operations. Specialized for your hardware. They run entirely locally in your browser. Now available on Hugging Face 🤗 Tell your agent to use @huggingface/kernels in your next project.
15
63
473
21,024
Sorry, absolute BANGER alert?!
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
2
8
571