It’s been great working on DiG-bench with @jcrwhittington and everyone else: a new text-based benchmark for in-context discovery (Discovery in Games). DiG-bench is within-domain for language models, but contains new text-based concepts and mechanisms that have to be discovered by interacting within a game. Players and models are told very little about the task or how it works, and have to figure it out for themselves. We show that while frontier models have improved a lot over the last few months, it's easy to make discovery games they can't beat but humans can. Although dig-bench is interactive, some of the original inspiration came from copy-cat, Doug Hofstadter and @MelMitchell1's early test of abstraction and analogy. @cocosci_lab (@Princeton) @MITCoCoSci (@MIT) @SchmidhuberAI (@KAUST_News) @misovalko (@Inria) @tri_dao (@PrincetonCS) @RMBattleday @zebkDotCom @FraserGreenlee @akaijsa @ClareMaguire @TimMuller1 @kubicek_ales @physicscat0x7d @SukritSumant @thoughtchannel_
We’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our novel discovery games and seeing how they score. Each game is a text-based environment, so they probe discovery capabilities in the natural domain of language models, rather than requiring additional, potentially confounding, visual understanding. TL;DR frontier models have improved a lot over the last few months. But they are still stumped by some surprisingly simple problems, even in their native text domain. With @cocosci_lab (@Princeton) @MITCoCoSci (@MIT) @SchmidhuberAI (@KAUST_News) @misovalko (@Inria) @tri_dao (@PrincetonCS) @RMBattleday @zebkDotCom @FraserGreenlee @akaijsa @ClareMaguire @TimMuller1 @kubicek_ales @physicscat0x7d @SukritSumant @thoughtchannel_ (1/5)
1
8
28
5,840
Ruairidh Battleday retweeted
in light of the landmark Meta settlement today, resurfacing this from March. Meta is set to pay up to $17B AND fundamentally change how Instagram and Facebook work for teens: time limits, nighttime blocks, school-hour notification restrictions, age assurance, limits on likes + beauty filters, and more. the era of “just add parental controls” is ending. how we design technology for young people matters. plz take a read ↓
1
4
10
4,434
Ruairidh Battleday retweeted
Our group discovered that reasoning models produce fractals when asked to solve hard problems. We can use nonlinear dynamics to probe the thinking processes of recurrent depth models on Sudoku, mathematics, and even ARC-AGI (1/N) arxiv.org/abs/2609.04963
112
624
4,708
455,379
Ruairidh Battleday retweeted
🎙️Full tech report on how to actually achieve SovereignAI, from data to model training, values, infrastructure, to large-scale deep research. Bridging up to 7 months of Frontier AI development with a $450k training Continual Learning run 🧠 📜 Paper: tinyurl.com/yk347xb 💻Model: tinyurl.com/2ju6z7dc Partners: @imperialcollege, @datologyai, @LambdaAPI
14
31
149
718,383
Ruairidh Battleday retweeted
very proud to have worked on the world's best open-weight VLA - there's a lot of fascinating parts of this tech report that I will call out in separate tweets. but the combination of the null-expert MoE, the joint VLM data mixture, and scaling laws for transfer learning from unlabeled video data make this a very groundbreaking release. The weights are out now on HF!! try it out
Today we're releasing Isaac 0.5: 36B dynamic MoE, open weight 🤗 embodied foundation model. Isaac combines multimodal video understanding, embodied reasoning and robot control into a single, sparse backbone.
7
7
116
8,986
Ruairidh Battleday retweeted
New #preprint reviewing and conceptually mapping the major theories in the field of aging (@LPiolopez, Navneet Jawanda) "Theories of Aging: From Damage and Programmed Theories to Goal-Directedness" preprints.org/manuscript/202… "Aging is an extensive biological process characterized by morphological and functional alterations at different biological scales, resulting in a systematic decline in biological functions ultimately leading to death. Overall, two main types of theories have been proposed: damage-based theories and programmatic theories. We propose a third option, framed in the cognitive perspective on multiscale biological systems. Here, we review the different theories of aging, organize them in a conceptual hierarchy, and integrate our new model within the existing frameworks: aging as a consequence of loss of goal-directedness (in anatomical space). In this model, aging is driven by a dynamical systems-level disruption of homeostatic alignment within cellular collectives that build and repair a healthy body during development and maturation. In this model, morphogenesis is a homeostatic (goal-seeking) process; left with no goals after the developmental phase has been completed, aging can occur in the absence of damage via the functional disbanding of body components from their original aligned goal state. We suggest a roadmap with important implications for regenerative medicine and aging."
38
88
440
76,198
Ruairidh Battleday retweeted
If we want AI to accelerate scientific discovery, models should be able to make new discoveries in the first place. But until now, we didn’t have many concrete ways to evaluate this ability. That’s why @jcrwhittington and our fellow @RMBattleday co-led the development of DiG-bench. This benchmark evaluates how well LLMs are able to beat text-based discovery games, where they have to discover the rules and the win condition. Some interesting findings: → Models have gotten a lot better over the past months, but they still struggle with some surprisingly simple problems → There’s a large gap in capabilities between closed-source model and open models → Discovering the rules of the games was the main bottleneck in beating them Check out their post for more details on how games are structured, how well specific models do, and to play the games yourself in your browser ↓
We’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our novel discovery games and seeing how they score. Each game is a text-based environment, so they probe discovery capabilities in the natural domain of language models, rather than requiring additional, potentially confounding, visual understanding. TL;DR frontier models have improved a lot over the last few months. But they are still stumped by some surprisingly simple problems, even in their native text domain. With @cocosci_lab (@Princeton) @MITCoCoSci (@MIT) @SchmidhuberAI (@KAUST_News) @misovalko (@Inria) @tri_dao (@PrincetonCS) @RMBattleday @zebkDotCom @FraserGreenlee @akaijsa @ClareMaguire @TimMuller1 @kubicek_ales @physicscat0x7d @SukritSumant @thoughtchannel_ (1/5)
2
2
4
515
Ruairidh Battleday retweeted
OUR BENCHMARK RELEASE! Give an agent the rules and it can often plan. Ask it to discover the rules, and things get interesting. Very interesting! The scientific object here is not planning alone. Each DiG-bench game is a latent-rule POMDP: the agent must identify hidden dynamics and a hidden objective while simultaneously controlling the system under finite lives and steps. Actions therefore have two jobs, changing the world and reducing uncertainty about it. The result I find most striking is Gemini’s mean game win rate: 93.5% with the rules, versus 14.9% without them. This suggests that the hard part is not only acting with a world model, but acquiring a useful one online. The next nice experiment would be scientist/controller split: let one agent explore and write a bounded theory, then give only that theory to a fresh agent solving unseen levels. That would separate transferable discovery from a lucky trajectory or trial and error. The 21 public games are live at digbench.ai. Fair warning: “text-only” does not mean easy. A blank exam sheet is also text-only. :) Great fun working with @RMBattleday, @jcrwhittington, @zebkDotCom, @akaijsa, Jimi Cullen-Drohan, Zihan Yan, @TimMuller1, @ClareMaguire, Seb Wilkes, @kubicek_ales, @FraserGreenlee, @SukritSumant, @physicscat0x7d, @SchmidhuberAI, Josh Tenenbaum of @mitbrainandcog and @cocosci_lab, together with @thoughtchannel_ and @Inria.
We’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our novel discovery games and seeing how they score. Each game is a text-based environment, so they probe discovery capabilities in the natural domain of language models, rather than requiring additional, potentially confounding, visual understanding. TL;DR frontier models have improved a lot over the last few months. But they are still stumped by some surprisingly simple problems, even in their native text domain. With @cocosci_lab (@Princeton) @MITCoCoSci (@MIT) @SchmidhuberAI (@KAUST_News) @misovalko (@Inria) @tri_dao (@PrincetonCS) @RMBattleday @zebkDotCom @FraserGreenlee @akaijsa @ClareMaguire @TimMuller1 @kubicek_ales @physicscat0x7d @SukritSumant @thoughtchannel_ (1/5)
2
6
35
2,070
Just seeing this post now! It was great to join everyone on this special issue on World Models (in the oldest journal in the world! Where lots of the seminal discoveries of the scientific revolution were published). We have a piece on AI for Science: The Easy vs Hard Problems.
It is the deepest honor to have been joined by Michael Levin (@drmichaellevin), Victoria Klimaj, Zahra Sheikhbahaee (@zah_bah), Dalton Sakthivadivel (@DaltonSakthi), Adeel Razi (@adeelrazi), David Ha (@hardmaru), Nick Hay, Kevin Schmidt, Irina Rish (@irinarish), David Krakauer (@sfiscience), Melanie Mitchell (@MelMitchell1), Samuel Gershman (@gershbrain), and Joshua Tenenbaum in organizing this special issue of the Royal Society’s (@RSocPublishing) Philosophical Transactions A: “World models, A(G)I, and the Hard problems of life-mind continuity: Toward a unified understanding of natural and artificial intelligence” royalsocietypublishing.org/r… This collection was motivated by a question with far reaching implications, ranging from the fundamental nature(s) of mind to choices that may determine the future of our civilization/species: what kinds of world modeling capabilities are likely to be realized by which kinds of minds and what world might we be in with respect to increasingly advanced artificial intelligences? Will the scaling and refinement of present approaches result in AI with human-like (and beyond) cognitive abilities, or do we need radically different paradigms that more closely follow the principles of natural intelligence? Learning “world models” to predict/compress information may be how biological learners so efficiently learn (to learn) to achieve goals and generalize that knowledge across a broad range of task environments. World models may also be useful for reverse-engineering forms of “System 2” cognition, or the self-reflexive, deliberate, multi-step reasoning associated with cognitive capabilities that may be unique to humans. Predictive models that reflect how the world may be causally modified by actions allow agents to adaptively control their behavior with flexibility and context-sensitivity. Spatiotemporally and causally coherent models of the physical world may not only be the key for creating AIs that we can rely on for real-world deployment, but may even be the (dynamic) core of conscious cognition. The contributions to this special issue consider the varieties of world models worth modeling from diverse points of view: Douglas Hofstadter explores whether sufficiently coherent self-referential world modeling could ground meaning, consciousness, and a genuine “I” in future AI systems. David Krakauer (@sfiscience), Melanie Mitchell (@MelMitchell1), and John Krakauer (@blamlab) examine the principles of emergent intelligence from a complex systems perspective. Alexander Ku (@alex_y_ku), Declan Campbell, Xuechunzi Bai (@baixuechunzi), Jiayi Geng (@JiayiiGeng), Ryan Liu (@theryanliu), Raja Marjieh (@RajaMarjieh), R. Thomas McCoy (@RTomMcCoy), Andrew Nam, Ilia Sucholutsky (@sucholutsky), Liyi Zhang (@LiyiZhang_Leo), Jian-Qiao Zhu (@JQ_Zhu), and Thomas Griffiths (@cocosci_lab) argue for using the tools of cognitive science to understand and evaluate LLMs across multiple levels of analysis. Evelina Leivada (@EvelinaLeivada), Gary Marcus (@GaryMarcus), Fritz Günther, and Elliot Murphy (@ElliotMurphy91) test whether LLMs deeply understand language and the “world behind words,” or primarily learn surface statistical regularities. Pedro Tsividis (@ptsividis), João Loula, Jake Burga, Juan Pablo Rodriguez, Sergio Arnaud, Nate Foss (@_npfoss), Andres Campero, Ajay Subramanian (@ajaysub110), Thomas Pouncy, Samuel Gershman (@gershbrain), and Joshua Tenenbaum introduce a theory-based meta-learning architecture inspired by the remarkable flexibility and efficiency of human cognition. Eunice Yiu (@eunice_yiu_), Kelsey Allen, Shiry Ginosar (@shiryginosar), and Alison Gopnik (@AlisonGopnik) explore empowerment, controllability, and causal reasoning as means of understanding the remarkable learning abilities of both child and adult minds. Nadav Amir, Stas Tiomkin, and Angela Langdon investigate how goals shape the structure of experience and how the world modeling abilities of natural intelligences may be inseparable from values. Vickram Premakumar, Michael Vaiana, Florin Pop (@FlorinPop17), Judd Rosenblatt (@juddrosenblatt), Diogo Schwerz de Lucena, Kirsten Ziman, and Michael Graziano show unexpected benefits of self-modeling as an inductive bias and regularizer for training artificial agents. Hanlin Zhu, Baihe Huang, and Stuart Russell analyze why model-based reinforcement learning may fundamentally outperform model-free approaches in representational efficiency. Bradly Alicea (@balicea1), Morgan Hough (@mhough), Amanda Nelson, and Jesse Parent (@JesParent) revisit fundamental cybernetic principles of regulation, adaptation, and world modeling across a wide assortment of complex adaptive systems. Francesco Sacco (@FrancescoSacco1), Dalton Sakthivadivel (@DaltonSakthi), and Michael Levin explore topological constraints on self-organization and suggest that biological systems maintain long-range coherence in ways that are fundamentally different from current transformer architectures. Georg Northoff (@NorthoffL), Yasir Catal, and Samira Abbasi examine how biological intelligence may depend on capabilities for flexible “inner time” to ensure adaptive alignment between the dynamics of system and world. Nicolas Rouleau (@DrNRouleau) and Michael Levin explore whether theories of consciousness generalize beyond brains to unconventional embodiments and living systems more broadly. Benjamin Lyons and Michael Levin investigate economies and collective intelligence as systems coordinated by “cognitive glues” in the form of shared models of scarcity and value. Katherine Collins (@katie_m_collins), Umang Bhatt (@umangsbhatt), and Ilia Sucholutsky (@sucholutsky) consider “Rogers’ paradox” to demonstrate ways in which collective learning is impacted by different kinds of human-AI interactions. Ruairidh Battleday (@RMBattleday) and Samuel Gershman (@gershbrain) distinguish between the “easy” and “hard” problems of science, and describe how while current AI systems demonstrate powerful narrow forms of optimization with respect to well-defined inference-spaces, further developments are needed for achieving capabilities for novel scientific discovery. Fritz Breithaupt (@FritzBreithaupt) explores narrative world models and the roles of uncertainty and transformative experiences in natural intelligences, suggesting that coherent agency may depend on better understanding human-like meaning-making. Taken together, these diverse perspectives suggest that while LLMs can clearly learn powerful generative models of language, they likely do so without having world models of sufficient spatiotemporal and causal coherence to achieve human-like reasoning abilities, capacities for generating subjective conscious experiences, or pathways to realizing artificial general superintelligence. However, by further developing world modeling architectures, we may eventually be able to create forms of intelligence that recapitulate the remarkable flexibility and generality of human intelligence. Finally, enhanced (e.g. more coherent/integrated) world models may not only afford expanded capabilities, but could potentially help ensure that increasingly powerful AI systems achieve both inner and outer alignment with human(e) values.
4
15
900
It was also very cool to be on a special issue with Doug Hofstadter, someone who I've admired for years and got me into thinking about thinking in the first place!
4
161
The next project at @thought_channel: 2 Weeks of Discovery 🎇
1
2
3
1,294
Come and see Prof. Karl Friston (@ucl) and @jcrwhittington (@UniofOxford) discuss the leading computational theories on the brain later this evening! A once-in-a-generation event 🔥 Link in comments below:
1
2
10
871
It looks as though we are going to have a full house at neuroMONSTER this year! Last week we hit our full capacity. It's inspiring to see so many people fascinated by the mystery of the brain and the computations underlying intelligence. @thought_channel
1
1
7
877
Really excited to launch our bi-monthly seminar series at @thought_channel! Hear new ideas about intelligence from leading researchers in AI, Neuroscience, and more ☄️
1
2
8
1,245
Thanks @smusslick for kicking us off!
2
173