pretraining lead @arcee_ai • opinions my own

🌖
really cool to see people pick nac up internally, and to see it go from a small project to this. i use this for everything right now, from running and babysitting experiments to patching bugs to more auto-research-shaped tasks. the team really polished it up, take a look!
We’re thrilled to introduce NAC, and a beta of our expanded Open Models API. NAC is an internal harness our research team has used daily since April. It’s built exclusively for long running, asynchronous and hands off work. A significant portion of all code committed to our pre-training, post-training and data pipelines in the last 3 months has been powered by NAC, often driven remotely from our phones. Started as an internal side project from @stochasticchasm, it was quickly adopted internally, and we’re happy to open it up to everyone.
8
8
83
12,146
been wanting to test this out for a long time and so so so happy that someone has done this. great work
Want to train your SWE agents 2× faster 🏎️💨? We introduce KV-streams, which, rather than flushing the KV cache on agentic compaction, we preserve it and keep generating. This avoids the re-prefill cost on both the inference engine and the trainer. Unlike other works that assume a fixed compaction strategy, we show that KV-streams can work with any compaction method in agentic tasks. Using what we learn from experiments in text-based games. We scale KV-streams to SWE tasks, where we match the performance of re-prefill compaction and full context in half the time 🚀 arxiv.org/pdf/2609.35750 🧵
2
5
70
8,111
current mood is
current mood is
24
1,886
roon where is my dot
5
43
2,166
here we go
4
260
stochasm retweeted
you think thats a faithfully verbalized chain of computation you're reading?
3
42
550
11,876
the models really do still make the silliest system design choices sometimes
4
62
2,039
>measured by the magnitude of the AdamW-preconditioned inner product between the gradient on a program’s output and the learner’s recent parameter movement. everyone's seen this by now, but i really think this part is so clean. is there any similar work for online data selection?
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
11
4
182
15,678
yooo HySparse2! i can't believe they also went YOCO. what's with the convergent evolution between dsv4.1 and this? i'll read the full paper later today probably
MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today. Less prefill, a smaller KV cache, better long-context retrieval—and we got all three at once. Compared with MiMo-V2.6's Hybrid SWA architecture: • 5.02× lower prefill FLOPs at 1M tokens • 4.5× smaller KV cache at 1M tokens • Better MRCRv2 and RULER-v2 scores, plus lower AgentPPL and LongPPL Why build a new architecture? Agentic inference is a very different workload. Each round, a short action can return a long observation that needs to be prefilled, while the context keeps growing. That puts prefill cost, KV-cache size, and retrieval accuracy on the critical path at the same time. HySparse2 tackles all three with two levels of KV sharing: • KV Bridging: Following YOCO, full-attention layers in the cross-decoder build their K/V from self-decoder hidden states. • KV Reuse: Within each hybrid block, sparse layers reuse the preceding full-attention layer's KV cache and selection indices. Two more changes: token-level selection replaces block-level selection, and a forced window of recent tokens replaces the separate SWA branch, so local and global tokens share one KV cache. Since all cross-decoder KV caches now come from the self-decoder, prefill can stop once the self-decoder finishes. Paper: arxiv.org/pdf/2609.26368
6
1
98
10,796
i didn't read the full paper today...
3
373
damn this is hella informative and fun to follow
in @stochasticchasm tradition, I am going to do a paper breakdown of the new DeepSeek Elastic Compute after reading through it. looks very nerd snipe-y and interested in the lessons they learned. arxiv.org/pdf/2609.22978
9
257
23,454
what are the most approachable ml compiler project codebases and what are the most modern
10
1
95
5,594
i thought this was a fun read, trying to figure out how much of a capability boost swarming gives us. a lot of swarming writeups coming out recently
Just dropped my new article. Hope you enjoy it :) scaling01.substack.com/p/acc…
1
23
2,363
>there's a lot of really impressive infra design that gets stuffed into a few sentences in the report there's many sandbox systems out now but glad deepseek wrote this up
Replying to @eliebakouch
my own final impressions is that as usual for a deepseek report, there's a lot of really impressive infra design that gets stuffed into a few sentences in the report. the standout highlight is how tiny the kv cache size is, it's honestly absurd. i really feel for how all this would have come together in a more unstable way than usual, it seems like there are far more moving parts than before, and nailing every little thing with so many moving parts is very very difficult. very fun read and i'm sure i'll be re-reading it tomorrow and i'll find something i missed
4
17
244
20,075
okay let's read this
Replying to @llm_wizard
Full tech report is so detailed, btw, @stochasticchasm - it's time.
4
9
218
13,714
this is probably my favorite out of the various infra tricks they've got. rollout scheduling can have such a outsized effect on inference efficiency
1
5
287
probably wrapping up here, but great paper. absolute aura farm to put the metrics in public
6
257