I like simple & minimal examples and creative ideas. Foundations/theory of deep learning, AI. GDM | Prev: PhD, CMU

New York, NY
In my next blogpost, I write about how I view technical communication: it's like trying to communicate an escape route to someone without a map but with a catch: you're not with them. You only have a walkie-talkie. Also, they're in panic.
2
8
43
5,280
Vaishnavh Nagarajan retweeted
We (@ChandarLab) are continuing our tradition of helping underrepresented students apply for grad school by reviewing their application materials and giving feedback! Please retweet this for maximum reach! Deadline Nov 1st!
🎓 Graduate Program Application Mentorship Applying to grad school doesn't have to be overwhelming! Get personalized 1:1 support from students who’ve successfully navigated the process. 🔗 Apply now: forms.gle/FyjEuLmjcdREF3Ba6 🗓️ Deadline: Nov 1, 2026 👇 Here is what you get:
13
72
7,489
Is no one gonna talk about the fact that the surname in reverse reads "No Slop"
44
188
2,849
123,062
At that point, it was just a cute observation that showed up when I was trying to characterize the space explored by SGD for a fixed random init (in the hope of understanding why deep networks generalize well).
Replying to @_arohan_
First linear mode connectivity (not named that until Frankle, Dziugaite, et al 2019) observation was by @_vaishnavh and it was on test loss when swapping the training data. LMC in Frankle, Dziugaite, et al 2019 was defined on train, IIUC, but we definitely observed it on test (and think it was in the abstract). There's been a lot of cool work since then.
3
36
3,632
Loved this analogy between how we perceive a fully-fledged scientific discovery standing on the shoulders of primitive glimpses of itself and how we perceive an orchestra. (From Loren Eiseley's The Firmament of Time.)
10
440
Vaishnavh Nagarajan retweeted
What if a foundation model could tell us how to modify its architecture to boost inference and reasoning instantaneously—without retraining? What if that tweak incurred near zero latency cost during generation and supported indefinite state tracking? arxiv.org/abs/2608.17981
22
101
538
132,400
Vaishnavh Nagarajan retweeted
First blog post: "Quantifying Pretraining Variance" Data and init seeds are familiar sources of run-to-run variance in LLM training. But floating point arithmetic order (e.g. from sharding differences) also affects things, and we find these effects are nearly as large!
7
11
119
6,435
stumbled upon this book and discovered a new love for science/science-adjacent writing that also has a poetic quality to it. seeking similar recommendations!
2
1
17
800
Vaishnavh Nagarajan retweeted
"How Transformers Learn to Plan via Multi-Token Prediction" has been accepted to #COLM2026! Why does MTP improve planning? We identify gradient decoupling that enables reverse reasoning: look to the goal, then trace the path back. 📄 arxiv.org/abs/2604.11912
44
251
14,007
Vaishnavh Nagarajan retweeted
Excited to share about what I’ve been working on over the past year: quantifying the capacity of a neuron😅 This led to a mathematical framework we call HOPE, which lets us rigorously deconstruct what deep networks might have learned. Paper: arxiv.org/abs/2607.21366
27
138
979
164,923
Such a simple, fundamental, well-posed and well-motivated question!
A question on synthetic data generation: If we want a language model to solve k-step arithmetic problems (such as a+b*c-d=?), with operands from 1 to 100, which training distribution should we use? A. Uniform distribution: Sample these k operands uniformly from 1 to 100 B. Power law: randomly shuffle 1-100 and impose an artificial power law. Sample these k operands according to this power law. ⚡Our ICML 2026 (spotlight) paper shows: Option B is better! Surprisingly, the same idea extends far beyond this simple example to many reasoning tasks that require implicit composition of multiple atomic skills, including multi-hop QAs and synthetic GSM problems. 📄Paper: arxiv.org/abs/2604.22951 📝Blog: zixuan-wang-dlt.github.io/po…
2
29
7,230
Some beliefs we may hold purely based off of gut feelings and instincts. It's important to recognize what those are and to try and reverse-engineer where those beliefs come from. If a rationale can be constructed for/against a belief, then we know what to do with the belief.
1
1
479
If not, we know that the belief is gut feeling and vibes (which is fine, but it's important to know that it's what it is).
141
Vaishnavh Nagarajan retweeted
Updated our paper on the foundations of memory in sequence models (with fresh insights, clearer writing and ablations). Our paper contrasts two distinct ways in which language models memorize and formulates the questions that arise from this. Will be presented at #ICML.
3
14
116
12,973
Vaishnavh Nagarajan retweeted
Go chat with Jacob to at ICML to hear how we’re rethinking pretraining with downstream behavior in mind: reducing catastrophic forgetting, improving adaptability, and preventing diversity collapse.
I'm in Korea at ICML this week and excited to present 4 papers on pretraining! theme is "how to do pretraining for continual learning" (short thread)
5
44
5,737
If you're looking for questions no one has studied before in the science of deep learning (specifically, about memory & geometry of representations in language models) find @ShNoroozi, @ElanRosenfeld and me tomo at <insert #ICML poster session here>
1/ We found that deep sequence models memorize atomic facts "geometrically" -- not as an associative lookup table as often imagined. This opens up practical questions on reasoning/memory/discovery, and also poses a theoretical "memorization puzzle."
2
8
53
4,643
Wed, Jul 8, 2026 5:00 PM – 6:45 PM KST HALL A #1611
263
I like the world map setting. Offers various ways to intervene and analyze the model's memory/representations (and makes it fun). Also, I'm a fan of curating the training task to illustrate something about the training dynamics and learning biases of models.
🚨 New Paper! (Part 2: Fine-Tuning) As seen in part 1, multi-task pretraining gives models clean, shared world representations. Now, say the world acquired a new set of cities: "Atlantis" cities. Let's fine-tune them into the network. 1/n
12
3,045
It's tempting to think that the lottery ticket hypothesis requires an exponentially large network for a subnetwork to win the lottery. This is incorrect--a small network suffices! We tend to imagine each subnetwork as throwing a dart in high-dim space to match a target
4
2
29
2,583
but what's played is *multiple*, *independent*, easier games of dart-throwing in multiple low-dims; the best dart from each game is put together to make the winning subnetwork. So you don't need exponentially many darts! I've elaborated on this here: vaishnavh.github.io/blog/lth…
1
10
729