Working on generalization at Q Labs. qlabs.sh/ Previously perception @nuro, math @caltech

SF
Akshay retweeted
We've figured out how to *pretrain transformers* with zeroth-order optimization and no backprop. Many of the core assumptions in optimization research are completely wrong. (paper out soon)
166
152
2,681
489,101
Scaling computational depth is super promising:
We think computational depth is the missing scaling axis, i.e. we should be doing a lot more deep learning! Every other axis has been scaled by OOMs over the past few years (params, data, sparsity, test-time reasoning), but depth has been stuck at ~100 layers since GPT-3. We've found that LLMs are both *severely* depth-bottlenecked and bad at using the depth they have, and that architectural interventions that lift this bottleneck efficiently lead to gains that increase with compute. w/ @akshayvegesna
3
2
46
4,273
Model architecture can improve scaling exponents!
We've been obsessed with bending the scaling laws recently and our new paper shows that looping with model growth improves the scaling exponent, leading to gain that compounds with compute! Everyone assumes architectural changes only give constant factor gains and pretraining progress mostly comes from data (e.g. @dwarkesh_sp's recent post). We found that model growth, looping, and boundary operators result in compute multipliers over standard transformers that grow exponentially with each OOM of compute. - 1.55x at 1e20 FLOPs and 2.7x projected at 1e25. - Matches GPT-3 13B on CORE with 20x less compute w/ @charllechen, @akshayvegesna, @andrewgwils 🧵
2
19
1,854
Akshay retweeted
a veryyy deep paper coming out soon
7
3
73
5,382
Akshay retweeted
this thursday we'll have jesse hoogland (@jesse_hoogland) to discuss singular learning theory (SLT). SLT establishes a connection between the geometry of the loss landscape and internal structure in models, offering a principled answer to why neural networks generalize. arxiv.org/pdf/2510.12077
3
5
52
14,753
Akshay retweeted
this thursday we'll have surya ganguli (@SuryaGanguli) to discuss the origins of scaling laws. his paper predicts scaling exponents of LLMs from measurable statistics of natural language. more generally, his work draws on ideas from statistical physics to understand deep learning. personally, i think scaling laws are an incredible phenomenon that needs serious study, and understanding their origins is crucial for improving them. arxiv.org/abs/2602.07488
10
20
159
26,526
Akshay retweeted
we'll have @AlexiGlad this thursday to discuss energy-based transformers. energy-based models (EBMs) are a more general framework for generative modeling than autoregression and diffusion (which is an implicit EBM). EBTs learn an explicit energy function for transformers, which allows test-time optimization over a learned landscape for stronger OOD generalization. register: luma.com/34g58cin arxiv.org/abs/2507.02092
3
30
9,039
Akshay retweeted
we recently published a paper on how to make the most of a multi-epoch budget when pretraining language models on limited data. here's a blog post about it: bishwasmandal.com/blog/q0-hy…
1
7
20
1,921
Dwarkesh's take is that RSI is automated AI research aimed at the sample efficiency problem I agree with this take, but human AI research in the sample efficiency space is surprisingly empty. It seems that people simply think that it's way too hard to make progress, which I'm pretty sure is false. It does require some new ideas though
New blog post: on the million-x sample efficiency gap between AIs and humans, and whether it matters: "The reason it is relatively easy for open source and previous laggards to catch up to within months of the frontier is that data is the real driver of progress. And data can be easily distilled from public APIs, whereas hyper-parameters and training tricks and architectural micro-optimizations cannot - if the latter were driving most of progress, then catching up would be harder than we are observing it to be. It is easy to forget how much data these models are trained on, and how much more it is than what we humans see in our lifetimes. We see these AIs as a galaxy glittering with capabilities, but at their center, invisible to the naked eye, holding all the constellations together, is an unimaginably massive black hole of data." Post in link below
1
13
2,390
Akshay retweeted
this thursday evening we'll have andrew gordon wilson (@andrewgwils) over to discuss epiplexity, compression, and understanding generalization in LLMs. it'll be a small group, in-person, hosted at our office in sf. dm me to join.
2
3
51
11,281
Akshay retweeted
hey everyone, i'm starting a tiny in-person reading group in SF for people working on generalization / science of deep learning. fundamentals matter more than ever and i think this'll become the biggest research area soon. we'll discuss big new ideas and bring in the best generalization researchers for talks. dm me if interested.
12
7
125
10,711
Akshay retweeted
1/ Now that we're running out of data, how do you optimally scale multi-epoch pretraining to hundreds of epochs? Our first paper from Q! q0 trains a population of models, instead of single model that saturates fast, reaching a dramatically lower loss at *every* epoch budget. w/ @bishmdl76 @akshayvegesna @ShmuelBerman
16
54
265
31,203
Cool presenting on why generalization in neural nets is less of a mystery than many make it out to be:
Last week we hosted the first ever YC Paper Club in Mountain View. We brought together great AI researchers and founders to discuss both the state of the art and what it actually takes to get it into production. Thanks to the following presenters: 0:12 - Intro from YC Visiting Partner @FrancoisChauba1 3:49 - Tanishq Kumar (@tanishqkumar07) — Speculative Speculative Decoding (arxiv.org/abs/2603.03251) 18:33 - Guangyao (Stannis) Zhou (@zhouguangyao) — Diffusion-MPC (arxiv.org/abs/2410.05364) 30:26 - Isaac Ward — LeWorldModeling (arxiv.org/abs/2603.19312) 43:54 - Akshay Vegesna (@akshayvegesna) — Deep Learning is Not So Mysterious or Different (arxiv.org/abs/2503.02113) 51:24 - Konwoo Kim (@konwookim) — Pretraining Under Infinite Compute (arxiv.org/pdf/2509.14786)
1
5
31
14,932
Akshay retweeted
it's encouraging to see recent momentum on theory of learning/generalization. i've been obsessed with this lately, few thoughts: - ml researchers often say we've been trying to solve generalization for decades. i'd argue we've solved everything except generalization. most of ML has been driven by empirical tweaks, and even most theoretical work (scaling laws, muP) is great science of neural nets as complex systems but not a theory of generalization. - most current attempts have the wrong framing. there's a lot of rigorous math that proves clean results but on toy problems. that's the wrong instinct for complex systems (a lesson from mech interp). you can't shortcut neural nets with proofs. at least initially, we need qualitative hand-wavy theories that empirically work. - a theory of generalization should work at the scale of frontier overparameterized LLMs and most importantly be predictive of how to train dramatically more data- and compute-efficient models, not just posthoc fit the deep learning we already have.
3
5
72
6,141
Akshay retweeted
it's really interesting how as a field matures the fundamental theories become dumber and dumber. in physics, we had differential equations and smooth manifolds, and people like wolfram showed such observed continuities emerge from much simpler discrete processes. in ML, we have fancy hessians, momentum, convergence theorems. but now we're starting to see discrete search work (alphaevolve, autoresearch). optimization works not because gradient descent is magic, but because the model is high dimensional and some of the directions end up working over enough steps. it's just guess and check scaled up. (image: wolfram)
13
14
301
42,089
If neolabs are successful and someone figures out a much better learning algorithm than transformers trained with backprop, it will prob be much less compute efficient to start out with. The current batch of chips are all built for fast matmul's. I think we'd have to be really lucky for the better learning algorithms to fit exactly this. Likely we will have to shoehorn it into existing chips to start and ultimately tweak the chips to make the models as compute efficient as possible.
2
9
480
Kinda gives “When you have eliminated all which is impossible, then whatever remains, however improbable, must be the truth." vibes on how to get human-like sample efficiency. My own view is getting 10,000x sample efficiency from loss function or simple architecture tweaks is prob impossible, so there's no choice but to replace gradient descent with new learning algorithms.
There's a quadrillion-dollar question at the heart of AI: Why are humans so much more sample efficient compared to LLM? There are three possible answers: 1. Architecture and hyperparameters (aka transformer vs whatever ‘algo’ cortical columns are implementing) 2. Learning rule (backprop vs whatever brain is doing) 3. Reward function @AdamMarblestone believes the answer is the reward function. ML likes to use pretty simple loss functions, like cross-entropy. These are easy to work with. But they might be too simple for sample-efficient learning. Adam thinks that, in humans, the large number of highly specialised cells in the ‘lizard brain’ might actually be encoding information for sophisticated loss functions, used for ‘training’ in the more sophisticated areas like the cortex and amygdala. Like: the human genome is barely 3 gigabytes (compare that to the TBs of parameters that encode frontier LLM weights). So how can it include all the information necessary to build highly intelligent learners? Well, if the key to sample-efficient learning resides in the loss function, even very complicated loss functions can still be expressed in a couple hundred lines of Python code.
2
6
644
Akshay retweeted
this is exactly what Slowrun aims to solve. i think the main problem is not the loss function but the learning algorithm: - gradient descent massively underutilizes compute relative to the data it's trained on. i think brains use a lot of compute to search for solutions that generalize, rather than doing computationally cheap but unexploratory curve fitting. - we've already gotten 10x better sample efficiency with Slowrun by pulling on two levers that let us utilize a lot of compute: heavy overparameterization (at least 3600x chinchilla) and multi-epoch training with regularization and data augmentation. - imo no single architecture or loss function change is going to solve sample efficiency. there's no replacement for spending compute on search. - humans aren't the ceiling either. current algos and humans are both very far from what's theoretically possible.
There's a quadrillion-dollar question at the heart of AI: Why are humans so much more sample efficient compared to LLM? There are three possible answers: 1. Architecture and hyperparameters (aka transformer vs whatever ‘algo’ cortical columns are implementing) 2. Learning rule (backprop vs whatever brain is doing) 3. Reward function @AdamMarblestone believes the answer is the reward function. ML likes to use pretty simple loss functions, like cross-entropy. These are easy to work with. But they might be too simple for sample-efficient learning. Adam thinks that, in humans, the large number of highly specialised cells in the ‘lizard brain’ might actually be encoding information for sophisticated loss functions, used for ‘training’ in the more sophisticated areas like the cortex and amygdala. Like: the human genome is barely 3 gigabytes (compare that to the TBs of parameters that encode frontier LLM weights). So how can it include all the information necessary to build highly intelligent learners? Well, if the key to sample-efficient learning resides in the loss function, even very complicated loss functions can still be expressed in a couple hundred lines of Python code.
2
5
53
5,600
Akshay retweeted
1/ nanogpt slowrun 🐢 update: we're focusing on occasional big data efficiency updates, but we had a lot of interesting additions in the last few weeks, here's the rundown: · multi-token prediction (@clark_kev) · looped transformers (@cs_serdar) · test-time training (TTT) (@Sam_Acqua) · stochastic logits averaging (@bishmdl76, @ShmuelBerman) · stochastic depth (@ChinmayKak) · IHA attention (@madhavsinghal_) · tuning for gradients norm (@zhiweixux) · probability averaging + scaling ensembles (@lzchen_ut) · MuonEq-R (@clark_kev)
1
17
116
10,745
Akshay retweeted
quick writeup on why i think diffusion isn't more data efficient than AR, since it seemed to surprise a lot of people: - the case for diffusion > AR ([1], [2]) rests on AR saturating at <5 epochs while diffusion can be trained for hundreds of epochs without overfitting. but that's AR with default regularization. with Slowrun we train AR for >30 epochs without overfitting using heavy regularization (15x standard weight decay and dropout), which captures the gains diffusion gets over hundreds of epochs. you can't push reg this hard on diffusion, the objective is already effectively regularizing the network - data augmentation is another lever that helps AR models: sequence permutation and token masking close a lot of the gap even without heavy regularization - [3] verifies this cleanly: simple dropout, weight decay, and token masking were enough to bridge the gap and even *surpass* diffusion. aligns with what we've seen [1] arxiv.org/abs/2511.03276 [2] arxiv.org/abs/2507.15857 [3] arxiv.org/abs/2510.04071
5
11
151
9,473