Since my previous post got some attention, let me use the opportunity for something more positive: to explain why Antoine Maillard’s work on ellipsoid fitting matters so much to me personally.
Cargèse, 2023. A conference I had organised with
@zdeborova A blackboard, and d²/4. That is where I first learned about this work from Antoine. I still have these photos from the lecture! You can see the threshold on the board. What you cannot see is everything those ideas would help set in motion.
Looking back, a surprising amount of what I now understand about neural networks, the relation between spectra and generalization, feature learning, emergence, power-law scaling, traces back to this line of work.
Because its importance was never just the number 1/4. The important thing was the IDEA. The statistical-physics analysis of Antoine Maillard and Dmitriy Kunisky predicted the threshold and the geometry of the solutions.
The work of Afonso Bandeira and Antoine Maillard made a crucial part of the programme rigorous, establishing the sharp threshold for approximate fitting through Gaussian equivalence.
arxiv.org/abs/2310.01169
arxiv.org/abs/2310.05787
For me, the powerful lesson was not simply “here is the answer to this problem.”
It was “here is a way to attack other problems.” And we did.
The next step in our own story came in 2024, when I invited Simon Martin to EPFL to tell us about his work with Francis
@BachFrancis . That visit helped bring these ideas together around a different question: what can we actually understand about learning in a two-layer neural network?
Together with Antoine Maillard, Emanuele Troiani, Simon and Lenka Zdeborová, we asked what happens when one learns a two-layer neural network with quadratic activation, whose width grows proportionally to the input dimension, from quadratically many samples.
And the key technical observation was precisely a connection to extensive-rank matrix denoising AND TO THE ELLIPSOID-FITTING PROBLEM.
We derived the asymptotic Bayes-optimal learning curve and introduced GAMP-RIE, an algorithm combining approximate message passing with rotationally invariant matrix denoising to achieve that performance.
This is the kind of thing we like here in Lausanne :-)
Maillard, Troiani, Martin, Krzakala & Zdeborová:
“Bayes-optimal learning of an extensive-width neural network from quadratically many samples”
NeurIPS 2024
arxiv.org/abs/2408.03733
Then Yizhou Xu, Antoine, Lenka and I pushed this further, putting the statistical-physics predictions on a rigorous footing.
We developed rigorous asymptotics and universality results for a broad class of structured matrix-sensing problems, including matrices whose rank grows proportionally to their dimension. Among the applications: establishing Bayes-optimal learning predictions for extensive-width quadratic neural networks.
Xu, Maillard, Zdeborová & Krzakala:
“Fundamental Limits of Matrix Sensing: Exact Asymptotics, Universality, and Applications”
COLT 2025
arxiv.org/abs/2503.14121
And then came one of the most satisfying consequences of this whole programme.
With Vittorio Erba, Emanuele Troiani and Lenka, we studied empirical risk minimization in overparameterized quadratic networks. L2 regularization on the network weights becomes NUCLEAR-NORM regularization in an equivalent convex matrix-sensing problem.
The mapping itself was not new. What excited us was bringing statistical-physics ideas and AMP to bear on it to derive sharp predictions.
Suddenly, we could connect global minima, generalization, memorization, weight spectra and the low-rank bias induced by weight decay.
This was exciting!!!
And d²/4 CAME BACK—as the interpolation threshold in the noise-dominated limit! More generally, the threshold depends on the target structure and the noise.
Erba, Troiani, Zdeborová & Krzakala:
“The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks”
NeurIPS 2025
arxiv.org/abs/2505.17958
But this was not the end of the story.
In a complementary direction, Simon Martin, Giulio Biroli and Francis Bach have analyzed the actual gradient-flow dynamics of extensive-width quadratic networks. They developed a high-dimensional dynamical description and characterized the resulting spectra, generalization and recovery thresholds.
Martin, Biroli & Bach:
“High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks”
JMLR 2026
arxiv.org/abs/2601.10483
Meanwhile, for us, the matrix/spectral viewpoint opened another door.
What if the target itself has a nontrivial spectrum? What if that spectrum follows a power law? How are different spectral modes learned as the amount of data increases?
Can this explain neural scaling laws? Can it give us a concrete mechanism for the emergence of newly learned features?
This led us to:
Defilippis, Xu, Girardin, Troiani, Erba, Zdeborová, Loureiro & Krzakala:
“Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime”
ICLR 2026
arxiv.org/abs/2509.24882
There, connections with matrix compressed sensing, LASSO and sparsity allowed us to derive phase diagrams for scaling exponents and connect generalization directly to the spectrum of the learned weights.
The spectral viewpoint makes the question concrete: which signal components are detectable, which remain unlearned, and how does that balance change as we get more data?
And remarkably, essentially the same viewpoint extends beyond quadratic neural networks.
In our work on a simplified single-head attention model, we again characterize learning through the spectrum of a learned matrix.
We obtain training and test errors, interpolation and recovery thresholds, and the full singular-value distribution of the learned query–key map, including low-rank structure and spectral outliers.
For power-law targets, learning proceeds through sequential spectral recovery: progressively weaker signal directions become learnable as the amount of data increases.
A concrete mechanism for emergence in this model—and scaling laws come out of the same analysis!
Boncoraglio, Erba, Troiani, Xu, Krzakala & Zdeborová:
“Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws”
ICML 2026
arxiv.org/abs/2509.24914
And this brings me back to why I care so much about the original story.
A good scientific paper does not merely produce an answer to one isolated problem.
Sometimes it gives a community a new connection, a new mapping, a new proof strategy.
A PATH.
For us, one path ran through:
Ellipsoid fitting → Gaussian equivalence → extensive-rank matrix estimation → quadratic neural networks → feature learning → spectra → scaling laws → attention.
Gaussian equivalence itself has a long, crazy history (don’t get me started on that one...)
Of course, this is not a single linear chain. Many researchers and ideas contributed at every stage, and several of these connections long predate the papers I have mentioned.
But having watched part of this story develop from very close by, I wanted to explain how Antoine’s work, with his collaborators, helped set this particular programme in motion—and how much we owe to it.
Looking at that blackboard again, I do not just see a threshold. I see ideas that helped us do research we did not know how to do before.
To me THAT is what “empowering researchers” actually looks like.