“You have to think like a mountain climber.” — Lasting Damage

Brooklyn
Neat!
The PhD thesis of my 16th PhD student, Fernando Hernandez Garcia, is now available. Title: Selective Reinitialization Algorithms for Preventing Plasticity Loss in Artificial Neural Networks Url: incompleteideas.net/papers/F… Abstract: In this dissertation, I study systems based on artificial neural networks that learn from nonstationary data. Learning from non-stationary data requires continual adaptation of the system, and is often referred to as continual learning. Developing systems capable of continual learning is a longstanding goal in artificial intelligence. In deep learning, the field concerned with designing and training deep neural networks, this goal remains elusive. This dissertation addresses a fundamental challenge in continual learning: loss of plasticity, a phenomenon where artificial neural network systems progressively lose their ability to learn from new data. While the phenomenon has been noted several times over the last three decades, it has remained understudied until recently. The work in this dissertation constitutes the first systematic demonstrations of plasticity loss, highlighting its persistence and importance. In this work, I provide a systematic demonstration of plasticity loss in a wide variety of deep learning systems. Across systems based on fully-connected networks, convolutional networks, residual networks, and vision transformers, plasticity degrades when learning continually. Notably, even systems employing normalization techniques, residual connections, and regularization—design choices that improve training stability—remain susceptible to plasticity loss. This evidence establishes that plasticity loss is pervasive and that deep learning systems trained with backpropagation are not suitable for continual learning. I explore the idea of selective reinitialization to prevent plasticity loss. This idea has been used in the past for improving generalization performance in deep learning systems. However, its use for preventing plasticity loss is a recent innovation pioneered by the continual backpropagation algorithm. The algorithm periodically sets new values for units in the network, making initialization a continuous process rather than a one-time operation. I demonstrate the effectiveness of continual backpropagation in preventing plasticity loss across a wide variety of settings. These demonstrations establish that plasticity loss, while pervasive, is not inherent to deep learning systems. I generalize continual backpropagation through an algorithm I call selective unit reinitialization. This general algorithm has three key components: a utility measure that ranks units by importance, a pruning criterion that selects which units to reinitialize, and a reinitialization method that assigns new values to the selected units. Continual backpropagation involves a specific choice of utility measure, pruning criterion, and reinitialization method. However, the general algorithm is not limited to the choices used in continual backpropagation. I present a study of how different choices for the three components of selective unit reinitialization affect its effectiveness at maintaining plasticity. This study establishes selective unit reinitialization as a general approach that can be tailored to each learning system to maintain plasticity. Finally, I propose a different approach to implementing selective reinitialization that operates at the weight level. I call the corresponding algorithm selective weight reinitialization. Reinitializing at the unit or weight level involves different trade-offs. Unit reinitialization minimally disrupts the network outputs, preserving stability during learning. However, the definition of a unit varies across network architectures, requiring additional engineering to make the approach effective. Weight reinitialization, in contrast, can be readily applied to any arbitrary network architecture, but can substantially affect network outputs, reducing training stability. The reduced learning stability can be remedied with L2 regularization, which stabilizes selective weight reinitialization but introduces an additional hyperparameter. Both selective unit and weight reinitialization successfully maintained plasticity across the systems tested, providing flexible approaches for different architectural and engineering constraints. Fernando is now a research scientist at Zyphra.
87
That is a neat trick. Nice to see useful applications of MCMC
SFT is not dead! 🥳 We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and forgetting less than RL and OPSD. 🤯 Following our prior work on reasoning with sampling, we now introduce sampling to the posttraining stack. 1/n
1
1
128
I was inspired by @kywch500's Tetris blog post from last year (kywch.github.io/blog/2025/12…) to do a bit of a replication. His W&B artifact hits about 550 lines cleared per game. Using shaped rewards, I varied model size and doing some "pretraining" on rollouts from a LP solver. It's fun to see what comes out. - Bad pretraining is really bad! Random initialization is better. - Pretraining helped model climb faster, but doesn't change final score - Bigger models climbed faster and reached higher scores - Hyperparams were optimized around 1x model size; using for 2x and 5x increased run variance
2
1
2
768
Me: Astra, please run these experiments and get the reward to X level. The experiments have these constraints which is what we want to test. Astra: Good news, I got it working! Here's what made the difference. Me: Super, now run this hyperpameter sweep Astra: Done! Results make sense! Me: Great, let's document approach using <framework> Astra: Will do! I will also clarify that only unsuccessful initial approaches used that experimental framework, and the working result relied on something that was clearly outside experimental constraints. Me: what
2
105
That'll teach me a lesson on when to review the exact code that's being run.
22
Researcher encounters politics 101. I've seen this multiple times before, and IMO is more a tell on the person complaining than anything else.
64
This looks great
NeurIPS / Claude rejected our paper "Life After Benchmark Saturation" but we continue to think it is an important contribution to evaluation science, so we hope you check it out here! arxiv.org/abs/2606.26158
1
157
These RL algorithms remind me a bit of doing neural net architecture design back in 2016. Nothing works, nothing works, nothing works... then it just takes off.
83
Oh, Astra. I love how quick you are to acknowledge reality
122
Michael Griffiths retweeted
Introducing Synthetic Hospital: an open, fully synthetic longitudinal EHR benchmark with verifiable ground truth! 1,268 patients, 5,602 encounters, zero PHI. Physicians could not reliably distinguish its charts from real ones. 📄 arxiv.org/abs/2609.30027 💻 github.com/sparkcpark/synthe… ✍️ sparkcpark.github.io/posts/f…
53
146
1,303
298,415
Michael Griffiths retweeted
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
84
460
3,551
737,900
I wish I'd read this tutorial (arxiv.org/abs/1805.00909) by Sergey Levine - co-founder of Physical Intelligence - a few years ago. It makes much clearer how RL algorithms relate to standard probabilistic inference (e.g. Bayesian modeling, etc). REINFORCE is especially nice, because you can get to it from updating a deployed model w/ "production data." Your ideal distribution (that maximizes rewards) is unknown, but you can get updates from an oracle. You can use the updates to "tilt" your current distribution and take an update step in that direction. You tilt with a weight from the oracle (the reward). Fun: if the oracle gives boolean rewards, this reduces to rejection sampling! You do a rollout, reject negatives and retain positives, and update based only on positives. That's just REINFORCE with no baseline adjustment. (Thanks to ChatGPT for the pretty image)
1
5
220
My hobby project du jour has been cloning @karpathy neat nanochat project in minimal-dependency Julia. We're about about 10% slower per gradient step, OK. So I was very surprised to see that when training a small 50M parameter model on a RTX 4090 we finished 30% faster (!). It turns out that nanochat had non-blocking data transfers disabled, so the Julia version benefited substantially. The PR turns out the be very small, which is nice! github.com/karpathy/nanochat…
2
1
6
787
Gosh, cross device is so hard. Either it's totally explicit and doing anything is tough, or it's implicit and you shoot yourself in thy foot. For each micro batch I have the line: ℒₜ += ℒₛ ... that accumulates number of tokens see so far for the loss sum so we can compute average. I accidentally set ℒₜ to be on CPU, so it was causing multiple waits per update. Crazy!
15
Are agents good agents for their principle?
➕➕ New research ➕➕ On the design of agents for markets and markets for agents. AI agents are increasingly entering markets, acting on behalf of people— How do we make sure they represent their people well, and that these markets are set up safely and fairly for everyone?
2
122