I post the papers I find interesting. There are so many papers published these days, and I frequently miss great papers. I appreciate paper recommendations via DM, but I tend to only post papers I discover on my own to keep my list personally curated.
6
3
116
17,398
Replying to @rosinality
LongCat has been using PID for its control for a while! I tried it for a while in Marin, but PIDs have far more hyperparams to tune than is desirable.
1
2
23
1,598
arxiv.org/abs/2609.39137 Aux loss free balancing -> Integral control Quantile balancing -> Proportional control This -> Integral-Derivative control.
5
15
113
19,595
arxiv.org/abs/2609.40295 Would it be better to filter out AI generated tokens for pretraining?
2
8
93
5,369
You can drop the vision encoder if pretraining compute is larger than 1e22 flops.
4
41
615
36,743
More efficient setup for the looped MoE. 2x experts, 0.5x looped layers, 2x loops, and attention untying. Now the looped transformer becomes a problem closer to better allocation of resources instead of inductive biases.
2
15
127
6,850
Rosinality retweeted
Added "Auto-character Coverage" to SentencePiece as a clean alternative to Byte-Level BPE (BBPE). It globally optimizes the vocabulary budget without invalid UTF-8 byte fragments, achieving comparable or better compression. google.github.io/sentencepie…
1
11
42
4,817
They are still avoiding using synthesized data, but maybe they have used data from more capable models? But for multimodal data they don't use synthetic data (which is the area where synthetic data is used extensively). They now explicitly mention the scaling ladder, their own crawling system.
As cited in the technical report it is a decoder-decoder architecture for KV cache compression (arxiv.org/abs/2405.05254). Very interesting.
1
23
2,689
As cited in the technical report it is a decoder-decoder architecture for KV cache compression (arxiv.org/abs/2405.05254). Very interesting.
Deepseek V4.1 Flash 552B total, 8/16B active with a new arch trained on 45T tokens, there are different active parameters for input/output tokens with the encoder/decoder arch, engram, new sparse attention, new mHC, native vision very high benchmarks (beating K3), insane efficiency, and as always amazing tech report this is probably the most novel arch i've seen in a while, pretty insane
2
9
64
9,255
I can't understand why people keep trying to say some company won the race after each model release. What is important for the model company is whether they have a roadmap and good direction and whether they are able to achieve it, not the model at each specific time point which eventually gets deprecated soon, and these things are generally hard to know from the outside of the company.
1
1
35
1,990
Maybe full-bandwidth transformer (not looped transformer) style architecture could allow more obscure cot? Though I think it will still be anchored around discrete tokens.
5
27
2,323
Great results when everyone talks about looped transformers. Looped transformers are now more compute-efficient compared to non-looped ones.
Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched? We introduce SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers. The answer is yes. And the advantage grows with scale. 🧵 1/8 Paper: arxiv.org/abs/2609.01343
1
5
42
5,685
Rosinality retweeted
Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched? We introduce SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers. The answer is yes. And the advantage grows with scale. 🧵 1/8 Paper: arxiv.org/abs/2609.01343
8
71
451
74,795
arxiv.org/abs/2608.30627 Inserting reasoning tokens into the pretraining data. This has been tried multiple times, but how scalable is it?
3
9
68
5,502
arxiv.org/abs/2608.24814 Effective learning rate, the ratio of learning rate and weight norm, governs training dynamics. This could allow transferring the settings across norm control methods (by matching effective learning rate).
3
18
166
10,819
arxiv.org/abs/2608.20061 Hyperparameter transfer attempt for 10T scale. Transfer over token horizon was done through a scaling law.
1
7
108
6,162
arxiv.org/abs/2608.19197 Synthetic environment generation through solver agent and environment generator dynamics. The rewards for environment generator are calculated using the difference of solver rewards conditioned on privileged information or not.
2
9
100
5,706
arxiv.org/abs/2608.18486 Introducing cross-layer connections is popular now. The problem is how to parallelize it (like Jacobi iterations) and whether it is enough for large-scale training.
13
113
6,600
arxiv.org/abs/2608.17981 Recurrent transformer that injects activations from the top layers to the bottom layers at the next step (arxiv.org/abs/2608.08888). Why does this work without training?
32
104
1,058
236,105
arxiv.org/abs/2608.17286 Scaling law estimation for text-to-image diffusion. One interesting result is that diffusion is forgiving for overtraining in the sense that the loss difference between compute optimal and overtrained models is relatively small.
8
71
4,641
arxiv.org/abs/2608.14071 Data repetition during pretraining, when non-repeated data is available to fill the remaining portion to keep TPP constant. It is another observation on how high quality data could be repeated more, with the twist that a larger model (!) and shorter LR decay tolerate more repetition better. This could interact with the "non-repeated" web data part, as it could be a balance of noise fitting between noisy unique data and high quality repeated data.
5
24
203
20,250