The PhD thesis of my 16th PhD student, Fernando Hernandez Garcia, is now available.
Title: Selective Reinitialization Algorithms for Preventing Plasticity Loss in Artificial Neural Networks
Url:
incompleteideas.net/papers/F…
Abstract:
In this dissertation, I study systems based on artificial neural networks that learn from nonstationary data. Learning from non-stationary data requires continual adaptation of the system, and is often referred to as continual learning. Developing systems capable of continual learning is a longstanding goal in artificial intelligence. In deep learning, the field concerned with designing and training deep neural networks, this goal remains elusive.
This dissertation addresses a fundamental challenge in continual learning: loss of plasticity, a phenomenon where artificial neural network systems progressively lose their ability to learn from new data. While the phenomenon has been noted several times over the last three decades, it has remained understudied until recently. The work in this dissertation constitutes the first systematic demonstrations of plasticity loss, highlighting its persistence and importance.
In this work, I provide a systematic demonstration of plasticity loss in a wide variety of deep learning systems. Across systems based on fully-connected networks, convolutional networks, residual networks, and vision transformers, plasticity degrades when learning continually. Notably, even systems employing normalization techniques, residual connections, and regularization—design choices that improve training stability—remain susceptible to plasticity loss. This evidence establishes that plasticity loss is pervasive and that deep learning systems trained with backpropagation are not suitable for continual learning.
I explore the idea of selective reinitialization to prevent plasticity loss. This idea has been used in the past for improving generalization performance in deep learning systems. However, its use for preventing plasticity loss is a recent innovation pioneered by the continual backpropagation algorithm. The algorithm periodically sets new values for units in the network, making initialization a continuous process rather than a one-time operation. I demonstrate the effectiveness of continual backpropagation in preventing plasticity loss across a wide variety of settings. These demonstrations establish that plasticity loss, while pervasive, is not inherent to deep learning systems.
I generalize continual backpropagation through an algorithm I call selective unit reinitialization. This general algorithm has three key components: a utility measure that ranks units by importance, a pruning criterion that selects which units to reinitialize, and a reinitialization method that assigns new values to the selected units. Continual backpropagation involves a specific choice of utility measure, pruning criterion, and reinitialization method. However, the general algorithm is not limited to the choices used in continual backpropagation. I present a study of how different choices for the three components of selective unit reinitialization affect its effectiveness at maintaining plasticity. This study establishes selective unit reinitialization as a general approach that can be tailored to each learning system to maintain plasticity.
Finally, I propose a different approach to implementing selective reinitialization that operates at the weight level. I call the corresponding algorithm selective weight reinitialization. Reinitializing at the unit or weight level involves different trade-offs. Unit reinitialization minimally disrupts the network outputs, preserving stability during learning. However, the definition of a unit varies across network architectures, requiring additional engineering to make the approach effective. Weight reinitialization, in contrast, can be readily applied to any arbitrary network architecture, but can substantially affect network outputs, reducing training stability. The reduced learning stability can be remedied with L2 regularization, which stabilizes selective weight reinitialization but introduces an additional hyperparameter. Both selective unit and weight reinitialization successfully maintained plasticity across the systems tested, providing flexible approaches for different architectural and engineering constraints.
Fernando is now a research scientist at Zyphra.