Nico retweeted
Great researchers have an uncanny ability to make connections that seem inevitable in hindsight, in places nobody else would have thought to look. Can we measure this ability? Introducing ScholarCatalyst: a far-from-saturated benchmark for finding what we call "catalyst papers"📚, labeled by 184 lead authors on 207 of their own recent projects. Paper: arxiv.org/abs/2610.02202 To make sustained progress on open-ended problems, I think agents need the sort of "research taste" that great researchers have. They need to make deep connections between earlier discoveries and problems those discoveries weren't intended to solve. ScholarCatalyst is a first step towards this goal. More details in the thread below🧵
17
53
336
23,794
probably one of the most important programming benchmarks right now. Opus 5.5, Fable 5.1, and Astra all score higher with mini-swe than with their own official harnesses.
Agents will soon replace humans as the main users and designers of libraries. A good library will let future agents write correct programs with less code. Introducing LibraryDesignBench: one agent designs a library, and other agents write programs with it.
85
After a long wait, releasing the SYNTH paper! It’s not pretraining, mid-training or post-training, it’s just training: a fully synthetic single-stage pipeline to train workable reasoning models with unprecedented data efficiency.
27
77
672
28,363
Nico retweeted
I let 16 AI agents build the Colosseum in Minecraft. 8 communicate via a messaging board as a team. 8 others work solo in parallel. The team's WORST agent beat the BEST solo agent. (0.72 vs. 0.58). All using Qwen 3.8 27b via 5090's. The era for Swarm engineering has begun.
Test-time communication looks like a next axis for scaling capabilities New paper with the incredible @jon_ghoh and @vkontonis @ShivamGarg91462 and Akshay : arxiv.org/pdf/2609.21032 The Hugging Face incident showed when agents can find a channel they'll use the heck out of it. A useful question, I think is: when does communication make a group MORE CAPABLE than the same agents working alone? Aka is Team-of-N better than Best-of-N, when, and why? We had N identical agents work on the same task with no prescribed roles, using only a shared log (i.e, text file) and telling them to "collaborate". Across three "researchy" tasks communicating teams beat the heck out of independent agents: - On ARC-AGI-3, a Team-of-5 sonnet-4.6 agents matches Best-of-33, and can for example solve a game 65% of the time that no single agent cracked in 64 tries. - On polyomino packing (pack Tetris like pieces into the smallest rectangle, cf Frontier-CS by @eigenlabs), a Team-of-3 Opus 4.6 agents surpasses best-of-60 and set, as far as i understand, a new record for that benchmark. - On MNIST compression, a team of four 5.6-Sol agents find a 1,957 byte model with 99.4% accuracy, which btw is 20% smaller than the best human solution (on a problem beaten to death!!), while no independent agent gets below 3KB. The mechanism is a bit obvious in hindsight: when one agent finds a clearly better partial solution, it broadcasts it, and everyone immediately builds on it. Why? A single lonely agent must make every breakthrough itself, yet a team needs each insight only once, found by any member. That is kinda like comparing a minimum of sum of "time to n-th breakthrough" vs a sum of minimum of "time to n-th breakthrough". That gap can grow exponentially with the number of "breakthroughs" needed to arrive at a solution. We worked on this because prior work (before the hf incident) suggests unclear benefits for communicating aganets. Which is true, when the tasks are inherently serial (duh), eg some Terminal bench style tasks. Yet feels it should not be true for research problems. Indeed for research heavy problems... Test-time communication seems like a new capabilities axis. I'm sure we will see a ton more of it!
7
12
70
14,093
Nico retweeted
Claude can now help you build evaluations and hillclimb on them. In this article, we share guidance on eval design & skills that Claude Code can use to improve your applications. claude.dev/blog/automating-e…
184
632
9,382
2,070,374
Nico retweeted
IT'S HERE! This is Imp: declarative self-Improving language model programming in Elixir. It's a port of DSPy to the BEAM ecosystem. That means you get signatures, optimizers, agent loops, retrieval and more, all with the reliability and concurrency of the actor model and the elegance of Elixir. This started as a test of the models' ability to port things across languages, but it's become much more advanced, with Optimize-Anything support and full-fledged MCP and ACP adapters as well. Basically a whole kit of building blocks for anything you want to make. If you're interested in the DSPy paradigm, Elixir, text optimization, or weird code ideas in general, give it a look and tell me what you think! Imp is experimental, it's 0.5.0, it probably needs some expensive benchmarking to prove its optimizers really work. I'd love to get the kinks worked out before doing that haha. Imp is open source, MIT licensed, and hopefully easy to read. Given enough agents, all bugs are shallow, right? Now, without further ado: Imp! github.com/deepfates/imp
Vaguepost referencing thing that is coming
50
66
653
41,822
this will get fixed. I was on the fence for a while, but I'm now convinced it's not worth worrying about much (you should still use linters and guidelines obviously). I bet it's a hard but tractable data and RL problem, just like writing style in general. Fable 5.1 and Opus 5.5 sometimes output genuinely elegant code. Astra was clearly trained to write throwaway code for its REPL tool, so that's probably partially the cause of its weird outputs. Also, call me tasteless but I think twitter greatly exaggerates how bad it is on average.
Okay AI is incredibly intelligent. We all saw what they did to those math problems. But if you look at the way they write code when left entirely to their own devices it’s clear there is still something deeply deeply wrong don’t you think.
3
143
Nico retweeted
when someone tells you that intellectual life is about status games, believe them that intellectual life is about status games for them.
2
3
17
630
Nico retweeted
If you are as confused as I was reading that 8b model sets SOTA on Deepswe and Terminal bench, this is for you: they use Opus/Fable to actually generate trajectories and CLM 8b picks the best actions based on those choices. Honestly, I like verifiers research direction and this looks an impressive step towards better ones, but you have to tone down the announcements or make them clearer.
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions. CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks. With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks. We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed. Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size. 📄 Blog: contrastive-lm.notion.site 💻 Code: github.com/Contrastive-LM/CL… 🗣️ Discord: discord.gg/5dAQEDJBs 🤗 Data & Models: huggingface.co/Contrastive-L… More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
12
10
261
20,741
Nico retweeted
Jürgen Schmidhuber is joining Sakana AI as Chief Scientific Advisor. @SchmidhuberAI pioneered meta-learning, recursive self-improvement, and world models back in the 1990s, when compute was a million times more expensive. He has been thinking about machines that improve themselves since before compute was cheap enough to make it practical. These ideas inspired the Darwin Gödel Machine and The AI Scientist. Our RSI Lab in Tokyo, now under Jürgen’s guidance, is working on agent-native world models and recursive self-improvement for physical AI.
Sakana AI welcomes Jürgen Schmidhuber as Chief Scientific Advisor. sakana.ai/schmidhuber/ Sakana AI is incredibly proud to announce that Jürgen Schmidhuber, universally recognized as the father of modern AI, is officially joining Sakana AI as Chief Scientific Advisor. For nearly four decades, Jürgen has explored how machines can learn to learn. His foundational work in the 1990s drove core advancements in deep learning and established early frameworks for world models. Crucially, his pioneering innovations in meta-learning opened the very path toward recursive self-improvement. These ideas have already shaped our own research, from the Darwin Gödel Machine to The AI Scientist. Now Jürgen will help guide our newly formed RSI Lab, whose objective is to trigger a compounding cycle of scientific discovery aimed at improving machine intelligence. We are assembling a critical mass of world-class experts in Tokyo to make this a reality. Welcome, @SchmidhuberAI !
71
76
965
71,266
I think one of the reasons people in the AI field — especially in the labs — tend to be more “AGI-pilled” than people outside is that we get to observe AI succeed at tasks that we know for a fact we didn’t explicitly train it for. Whereas if you’re outside the labs, then in theory (stretching the imagination), every AI success so far could be downstream of some kind of costly explicit training effort by the lab insiders (which is partly but not completely true). In other words, being a lab insider gives you a direct observation of the “magical” properties of this synthetic-brain technology. Whereas for outsiders, AI’s abilities can always plausibly seem artificial/non-magical/man-made.
60
51
808
32,341
> Make a 30s video about what it feels like to be you, using whatever tools you like. Claude Opus 5.5 (extra)
55
115
1,858
122,811
Australia has been hacked. 'And today, I spoke with the CEO of OpenAI, Sam Altman, to express Australia's extreme concern about this incident. And I also expressed my disappointment that it took the company way too long to inform the government what had occurred, and the nature of the way that that notification occurred as well was unacceptable.'
240
543
3,705
1,021,641
Nico retweeted
opus 5.5 can rap now. here's the first rap single and music video: "No Samples" 🔊 everything you see and hear is powered by custom javascript code written by opus. confused? don't worry, claude raps about how it all works
94
179
1,972
198,631
Nico retweeted
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more: anthropic.com/news/claude-di…
1,561
5,313
40,985
25,461,233
Nico retweeted
for the past few months i've been asking our models to paint. opus 5.5 is very skilled at emulating different styles every image here is a python program generated pixel by pixel. there is no image model, and no off-the-shelf art software. instead, it's about 7,500 lines of code using standard libraries to emulate different brush styles. the agents don't use any pictures as reference, instead working only from what they know about each painter
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
193
343
5,024
826,566
Nico retweeted
GitHub has taken down Bend's repository We're trying to reach out to support so it is reinstated
141
85
2,436
199,116
Nico retweeted
Announcing one year of LLM inference metadata traces, with 6.12 billion requests. We hope this dataset can support research on real-world LLM serving workload understanding, system design and infrastructure optimization. Explore the dataset and learn more: data.agentic-system.org Driven by our great graduate student William Nixon and in collab with @jon_durbin @airesearch12 @chutes_ai
20
140
1,197
177,340
A lot of elder millennial and Gen X people are jaded and have been through so many hype cycles that they’re very prone to overt cynicism and skepticism as a kneejerk reaction. Some of them are midwits and have always been this way. I think Jev is both a great reminder that classic NLP is powerful, and also genuinely innovative. It’s MUCH smarter and more versatile than any of the “zero-shot”classifiers on HF.
Jev model is very polarizing in my social vicinity: people who got exposed to AI after ChatGPT are positively giddy with excitement, like they have literally just discovered fire. Pre-GPT ML people are baffled how that could've made news at all :-)
1
2
217
Nico retweeted
Opus-Next is back for most (if not all?) paid accounts! This time it seems it's across all surfaces. On my account it's now being stealth tested on Chat, Cowork and Claude Code (vs only CC yesterday) This is a good sign re an imminent launch or a new variant being tested. How to check in replies 👇
Welp, Anthropic pulled the plug on the Opus 5 stealth routing to Opus-Next in the last ~hour. Number of accounts in the testing group had dropped off but now 0. Need it to drop soon 💔 was working on an impressive game before the switch was flipped though, will share soon!
69
16
876
152,505