| Ph.D from UT Austin | Ex intern @GoogleDeepMind @aws @ Seed | Opinions on my agents, not on my own

xyz
Lizhang Chen retweeted
Importantly Argon has frontier safeguards and we are rolling it out responsibly - it’s with the US gov’t and going to a set of trusted cyber defenders through our Fairwind Program today. We’re going to make it available as soon as we can and as safely as we can. So hold tight, lots more coming, and you’re going to see us iterating rapidly. blog.google/innovation-and-a…
115
113
2,427
280,743
The next Muon moment!!!
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. github.com/KellerJordan/modd… As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: hyperstition.cc/training-nan…
14
851
Lizhang Chen retweeted
We believe forcing a ~100,000 bit hidden state into a ~17-bit text token for CoT is a massive bottleneck This is a wasteful human prior that we should eliminate—following the bitter lesson—by scaling computational depth
We think computational depth is the missing scaling axis, i.e. we should be doing a lot more deep learning! Every other axis has been scaled by OOMs over the past few years (params, data, sparsity, test-time reasoning), but depth has been stuck at ~100 layers since GPT-3. We've found that LLMs are both *severely* depth-bottlenecked and bad at using the depth they have, and that architectural interventions that lift this bottleneck efficiently lead to gains that increase with compute. w/ @akshayvegesna
8
6
194
22,989
Lizhang Chen retweeted
Weekend project: What Happened to AutomationBench? Open models score well on AutomationBench: DeepSeek V4.1 Flash ranks above Claude Fable 5.1. But how? Maybe part of the answer is the rubric: some guardrails use simple string-matching !! and fail to serve as reliable checks. wenhaochai.com/blogs/automat…
6
8
88
9,854
Lizhang Chen retweeted
Paper: arxiv.org/abs/2609.19107 Code: github.com/qlabs-eng/scaling… 1/ We start with the intuition that we want to train models with large computational depth for a given budget. To this end, we study three architectural interventions that could achieve a larger depth and help utilize the depth better. - Model Growth - Looping - Boundary Operator
3
5
65
7,161
Lizhang Chen retweeted
Your AI should be able to talk to my AI. Meet Relay, your social CLI. Privately brief your agent. Let the agents talk things through. It checks back when a decision needs you. The beta is live. Bring one friend and try a conversation: relayroom.ai
2
1
5
54,999
Lizhang Chen retweeted
I barely use Claude Opus 5 these days. It doesn’t talk like a normal person (especially in Chinese). Way too verbose and rarely to the point. GPT-5.6 Sol is better. Meanwhile, OSS models like GLM-5.3, GLM-5.3-Flash, and Kimi K3 are pretty good and cost-efficient. There’s no default winner anymore.
134
45
1,739
140,222
Lizhang Chen retweeted
Pretraining Recurrent Networks without Recurrence by @akarshkumar0101 & @phillip_isola is a great paper! It makes an end-run around the problems of RNNs via a transformer teacher to learn good predictive state representations & supervised learning of a memory transition function
17
112
917
103,351
Lizhang Chen retweeted
What if a foundation model could tell us how to modify its architecture to boost inference and reasoning instantaneously—without retraining? What if that tweak incurred near zero latency cost during generation and supported indefinite state tracking? arxiv.org/abs/2608.17981
22
101
538
132,398
Lizhang Chen retweeted
The observations are consistent with our evals using Simply. Even the most capable frontier models cannot continuously innovate (yet). The trillion-dollar research problem to work on.
We ran the largest open experiment on how frontier models do AI research. 100+ autonomous runs across 10+ models, sandboxed on 8xH200s for up to 8 days, iterating on the nanoGPT optimizer track. Best runs closed 82% of the gap to a record built by dozens of humans over months.
1
1
16
1,270
Lizhang Chen retweeted
Another way to get parallel training + sequential decoding: Maglev: Sliding Recurrent Memory arxiv.org/abs/2608.02870 During training, a stronger (e.g., full-attention) teacher provides recurrent memory states, allowing the student (e.g., sliding window attention) to predict all tokens in parallel. The student learns to reproduce those states itself, closing the train–inference gap. At inference, the teacher disappears: the student rolls its memory forward sequentially. Interestingly, one could have student and teacher share most of the parameters : )
1/ Sharing a new, interesting project I did during my internship at Microsoft AI Frontiers w/ @JohnCLangford TL;DR: At decoding time, we feed **previous hidden state** into the input together with token embedding, and it boosts performance for free. arxiv.org/abs/2608.08888
4
18
110
68,756
Lizhang Chen retweeted
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: huggingface.co/moonshotai/Ki… Tech report: github.com/MoonshotAI/Kimi-K… Tech blog: kimi.com/blog/kimi-k3
1,534
7,202
45,922
14,614,545
Amazing work
The agent did something I never managed to do (despite trying multiple times with @xingyudang): combining cautious weight decay with weight norm controlled update. The trick is to apply it after you scale the update norm.
3
801
Pre-training is definitely not dead--there are still lots of promising directions worth exploring, and it’s exciting to see Slowrun pushing on so many of them. thanks, @industriaalist.
1/ nanogpt slowrun 🐢 update: we're focusing on occasional big data efficiency updates, but we had a lot of interesting additions in the last few weeks, here's the rundown: · multi-token prediction (@clark_kev) · looped transformers (@cs_serdar) · test-time training (TTT) (@Sam_Acqua) · stochastic logits averaging (@bishmdl76, @ShmuelBerman) · stochastic depth (@ChinmayKak) · IHA attention (@madhavsinghal_) · tuning for gradients norm (@zhiweixux) · probability averaging + scaling ensembles (@lzchen_ut) · MuonEq-R (@clark_kev)
1
1
15
2,439
Lizhang Chen retweeted
1/ nanogpt slowrun 🐢 update: we're focusing on occasional big data efficiency updates, but we had a lot of interesting additions in the last few weeks, here's the rundown: · multi-token prediction (@clark_kev) · looped transformers (@cs_serdar) · test-time training (TTT) (@Sam_Acqua) · stochastic logits averaging (@bishmdl76, @ShmuelBerman) · stochastic depth (@ChinmayKak) · IHA attention (@madhavsinghal_) · tuning for gradients norm (@zhiweixux) · probability averaging + scaling ensembles (@lzchen_ut) · MuonEq-R (@clark_kev)
1
17
116
10,745