now @ cursor, ex-cocosci @ mit, after thought ex-intern @ apple, google x

Brampton
It's finally here: Brampton Brampton is the world's most intelligent, creative, and fastest model. Brampton dramatically outperforms Grok 3, Claude 3.7 Sonnet, and GPT 4.5. Reply with "brampton" for early access.
2
595
Super excited to announce that I joined Cursor 🙃 Working on pretraining and longer term bets 🌞
60
10
1,153
38,192
Super cool work + also great to see ppl report the gen ppl frontiers :)
Self-conditioning can be distilled into one-step language generation. We show that self-conditioned flow LMs perform fixed-point iteration, and introduce fixed-point flow maps to compress both the iterations and the entire flow. At one step, FMLM★ achieves 112.5 gPPL at near-data entropy (5.37), substantially improving over FMLM’s 168.3 gPPL at 5.17 entropy. FMLM★ also achieves the best 2–4 step results. 📎 Paper: arxiv.org/abs/2607.00714 ⌨️ Code: github.com/Ugness/self-condi…
1
2
18
4,218
Sam Acquaviva retweeted
"Do not panic or hallucinate like a fragile Large Language Model."
3
9
157
5,588
So cool
How well can you describe the feature selectivity of a vision neuron … with words? Interpretability has long borrowed from neuroscience — and maybe it can give back too! 🧵
4
1,004
This is cool but disingenuous framing imo. All of the records were marginal, building on other records, so “out-outperformed all 1,016 other researchers” is a stretch lol I’m curious what this system’s performance would be if it couldn’t build on human records
OpenAI ran a hiring challenge, but the top candidate was one they couldn’t hire: our autonomous research agent, Aiden. In Parameter Golf, Aiden ran for 22 days, and out-outperformed all 1,016 other researchers: 🧵 (1/8)
2
16
2,541
Great work from Oscar on scaling up flow models!
Replying to @osclsd
🚨 Before concluding: As noted by @Sam_Acqua and many others, we all ought to be very skeptical of Gen PPL as a metric, especially in isolation. ❌ It is actually a bit crazy that we have been using it for so long. Hence, the additional metrics, and the presence of several qualitative samples in the appendix. Please have a look yourselves to get a better understanding of the sample quality! 🔍 There's comparisons across SD/non-SD, number of NFEs, and others.
1
4
1,050
Flow models are a promising alternative to autoregression. But the current standard of evaluating flow models is broken. The reported 3x improvement in 1024-step PPL since 2023 is closer to 1.1x if you control for sample entropy. (1/12)
7
28
166
50,749
One note: the main results here use 1024 samples while the main flow model results are << 1024 samples. I chose this to make comparison with diffusions easier and to make the point about the framework, not a given paper. I love flows.
1
1
7
962
Thanks to valuable discussions from @Chramblin, @nmboffi, @ReeceShuttle, @akshayvegesna, Samir, and other friends :)
1
6
844