The Stanford Artificial Intelligence Laboratory (SAIL), a leading #AI lab since 1963. ⛵️🤖 Emmy-winning video: piped.video/watch?v=Cn6nmWlu…

Stanford, CA
SWE-chat by @joabaum is a living dataset of human interactions with AI coding agents. 230K prompts from real developers collected in the wild are now available on HuggingFace 🚀 huggingface.co/datasets/SALT…
SWE-chat v2 is here! The largest dataset of coding agent interactions from real users in the wild has gotten even larger: 230K prompts from 18K sessions. SWE-chat has enabled incredible research (🧵) – excited to see what v2 unlocks for the community. Come find us at @COLM_conf!
6
3
31
8,184
Some interesting and perhaps unexpected properties of tokenization for language models!
Tokenizer Myths and Where to Find Them! Slides from talk last night - preprint forthcoming - feedback welcome 🙏 1/11
1
4
23
8,306
Stanford AI Lab retweeted
Most SWE benchmarks tell an agent exactly what to fix + where to look. Imagine if an agent could instead *proactively* surface and fix issues themselves (before a developer even runs into them!) We evaluate this capability at scale. Introducing SWE-sweep! Led by 🐐 @KLieret
Can LMs discover & fix bugs if you don't tell them what went wrong? To proactively maintain a repo, agents need to find problems before users do. SWE-sweep benchmarks this on 100 repos, 22 languages, 4k real bugs (numpy, php interpreter, lean kernel). Top models get <5%.
6
13
91
8,500
This work directly tackles the largely unexplored problem of submission overload at conferences. It argues that technical governance solutions can help everyone build trust and faith in our system, and in turn allows building on each other's contributions. Check it out!
🚨new method day! 🚨 Trust[ing results] in ML conferences is utterly broken. Let's fix it with an algorithm! We are excited to introduce 𝗧𝗿𝗮𝗶𝗻𝗶𝗻𝗴 𝗪𝗶𝘁𝗻𝗲𝘀𝘀𝗲𝘀, a system for “honesty” certificates of data and evaluations which a verifier can cheaply check and trust.
1
5
19
8,858
Stanford AI Lab retweeted
Bringing up the software stack for a new chip has always been a major bottleneck. That's about to change as we enter the era of AI as a compiler! The model directly translates source code (e.g., Triton) to low-level assembly (e.g., PTX), and a verifier checks functional correctness through symbolic / numeric analysis, as well as other properties such as race conditions and deadlocks! No need for manually constructed layers of IR and DSLs. Speedups over Triton baselines on B200: FlashAttention (1.37x) Mamba2 state forward (1.10x) Great work led by @cos_francois, in collaboration with Charly Castes and Thomas Bourgeat from EPFL! More to come soon!
7
18
196
21,862
Stanford AI Lab retweeted
Kinda crazy to get an Oral for your first paper - congratulations @alexxx_zzz6825 !! 🥳🎉🎊
Excited to share our EMNLP oral paper! We find that RL teaches model to traverse parametric knowledge more effectively, thus accessing knowledge previously inaccessible in instruction-tuned models. With @niloofar_mire, @RylanSchaeffer and Manasa Kaniselvan 🧵
3
4
28
7,576
Stanford AI Lab retweeted
Sorry for delay! Here is my work on zero order (ZO) optimization. We achieve SOTA for ZO methods on pretraining controlling for compute and parameters. ZO methods struggle to improve loss as model size grows because relative gradient variance increases linearly with the number of perturbed parameters, inhibiting large model training. All previous methods innovate on the optimizer, but adopt architectures designed for backprop. Instead, we design an architecture with ZO in mind, that caps the gradient variance as you increase model size. Introducing SOMA (Sharded Optimization Mixture of Assemblies). SOMA shards the model into tiny experts, and trains each expert independently, on disaggregated GPUs, without communicating gradients, activations or optimizer state during training. Inspired by the thalamus and cortical columns, each expert specializes on a subset of the train set based on a fixed router. At isocompute and isoparams, SOMA achieves lower loss vs. all monolithic ZO methods tested (e.g. EGGROLL, more perturbations w Vanilla SPSA, etc). ZO w SOMA continues to improve at larger model sizes. Additional inference benefits of this architecture, we can select top-k active experts over N trained experts to reduce inference compute by O(N/k), allowing us to tradeoff inference flops for accuracy wo retraining.
18
37
411
22,557
Stanford AI Lab retweeted
We found that LMs, large or small, text-only or multimodal, are susceptible to a shared problem -- 𝗼𝘃𝗲𝗿𝗰𝗼𝗻𝗳𝗶𝗱𝗲𝗻𝗰𝗲 𝗮𝗻𝗱 𝗼𝗯𝗹𝗶𝘃𝗶𝗼𝗻 𝘁𝗼 𝗱𝗲𝘁𝗮𝗶𝗹𝘀. When presented with a scenario that is uncommon with overwhelming pretraining data that encourages a shortcut, they take the shortcut confidently all the time. A picture that looks like a dog? Four legs. Location for the last Olympics? Paris. A question that looks like it needs detailed answers? Start a bullet point list. Simple finetuning won't cure this problem, especially when the baked-in behavior is so overwhelmingly represented in training -- but a simple modified preference learning technique would solve it perfectly. I'll be presenting this technique, Abductive Preference Learning, next Thursday at #COLM2026, come say hi! Joint work with Yijin Ni and @simon_ycl
2
2
22
9,334
Stanford AI Lab retweeted
Mercury Voice is here, bringing diffusion speed to agentic voice. Excited to see what developers build with it.
Today, we’re introducing Mercury Voice, a diffusion LLM specialized for agentic voice applications. Mercury Voice delivers 2x+ lower latency than models including GPT-6 Luna, Gemma 4 31B, and Claude Haiku 4.5, while beating them on a range of voice benchmarks. Enterprise customers interested in Mercury Voice can contact us at sales@inceptionlabs.ai to get access.
4
8
90
17,050
Stanford AI Lab retweeted
Looped transformer meets agent swarm. Our new @Neurips2026 paper shows how to loop compute across multiple agents. Different LLMs recursively share latent thoughts w/ each other. Better performance with up to 75% fewer tokens. Great work led by @Jiaru_Zou w/ awesome collaborators.
Excited to share that RecursiveMAS has been accepted at #NeurIPS2026! Amid discussions of GPT-6 Astra and RSI, Looped Transformers are drawing great attention for their ability to scale reasoning through recurrent depth in latent space. Looking beyond a single model, can we scale agent collaboration through recurrent depth in latent space as well? In RecursiveMAS, a team of agents reason just like a looped Transformer: each agent acts as a layer, passing and refining latent states across collaboration rounds. We pair this architecture with inner–outer loop training so the team learns to collaborate as a whole through recursion. We aim to make recurrent collaboration a new scaling axis for collective intelligence, combining stronger reasoning with faster inference. 💻 Project & Demo: recursivemas.github.io #RSI #LoopedTransformers #RecurrentDepth #LatentReasoning #MultiAgent #AgenticAI
1
41
216
23,445
Stanford AI Lab retweeted
Students sometimes ask me if it still makes sense, in this accelerating age, to pursue a PhD in AI. Perhaps counterintuitively, I think it's a great time to do so. I wrote up some thoughts on this here: web.mit.edu/phillipi/www/wri…
92
571
3,824
1,717,518
Stanford AI Lab retweeted
So excited to welcome @theworldlabs and @drfeifei to the @AMD family! I’ve always been a huge fan of Fei-Fei and her pioneering research in AI. Together, we’ll combine World Labs’ deep expertise in AI and world models with AMD’s compute leadership to power the future of AI and strengthen the open AI ecosystem. Can’t wait for all we’ll accomplish!
208
508
5,370
763,374
Stanford AI Lab retweeted
Excited to share our #NeurIPS2026 paper on benchmarking hallucinations in LLMs! We take a fresh world model-centric approach to measuring hallucinations in several environments including terminal, grid worlds, and chess. Thanks to the @DegenAI_Labs team for their hard work! Also grateful to collaborators at @Stanford, @StanfordAILab, @stanfordnlp, @CarnegieMellon, @LTIatCMU, @OhioState, and @OhioStateCSE.
Hallucinations are hard to measure because reality is messy. So we built our own world. Introducing HalluWorld 🌎, a controlled benchmark for LLM hallucinations. Now accepted to #NeurIPS2026 Evaluations & Datasets! 🧵 halluworld.ai
4
7
41
6,701
Stanford AI Lab retweeted
Congratulations to AIMI-affiliated faculty Ehsan Adeli, PhD, and colleagues on a new $25M @NIHAging grant to establish a Stanford center exploring how AI can help personalize care for people with Alzheimer’s disease and related dementias. 🔗 Learn more: stan.md/4hhURQA
1
1
11
3,384
Stanford AI Lab retweeted
Life update - I left OpenAI two months ago. I'm very thankful to have had the chance to contribute to agi and work with great teammates (miss ya'll ❤️). I'm now excited to start working on some of the problems I think will be important after agi...specifically robotics! 🤖🤖
70
17
1,089
93,671
Stanford AI Lab retweeted
PSA: While our names are similar, @mmitchell_ai and I are actually two completely different people 😂
Replying to @MelMitchell1
Because if you read the quote-thread you're continuing here, your stochastic parrot coauthor Timnit is literally calling it "the accurate description" of Astra and Fable. So of the two of you already wildly disagree on what the thing you both coined actually is nowadays, there's no hope the rest of us agrees.
12
6
253
56,919
Stanford AI Lab retweeted
1/10 Can we build a world model of human health? I'm thrilled to announce my first senior author paper: HealthFlux! Simulating how health evolves over a lifetime to find ways to prevent disease. doi.org/10.64898/2026.09.19.…
3
28
189
124,045
What would pretraining with zero real data look like? Can a randomly initialized model learn to generate all of its training data entirely through self-play? Find out below! :)
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
13
32
355
44,030
Stanford AI Lab retweeted
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions. CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks. With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks. We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed. Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size. 📄 Blog: contrastive-lm.notion.site 💻 Code: github.com/Contrastive-LM/CL… 🗣️ Discord: discord.gg/5dAQEDJBs 🤗 Data & Models: huggingface.co/Contrastive-L… More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
189
638
5,537
1,053,569