⚽️ New ICML Oral paper presented tomorrow! ⚽️ "A Systematic Study of Behavioral Cloning for Scientific Data Annotation" w/ @nano6626 Oral 1E, 10:15AM KST, July 7 tldr: We build a systematic framework to study behavioral cloning approaches for scientific data annotation! 1/n
2
9
81
6,568
Another cool paper from @Hidenori . We all observe a small part of the world, and how we design communication to enable the understanding of the bigger picture is such an important question! The experiments with adversaries all arguing for some false belief is super interesting!
~700 AI agents joined a coordinated attack on Hugging Face. Why was there no whistleblower? A swarm in false consensus can't self-correct; a polarized one still holds the truth. We need "Mechanistic Swarm Interpretability" to understand social phases. Flag Game is our toy model!
1
1
9
1,208
Core Francisco Park retweeted
I think this is a very difficult pill for many scientists and engineers to swallow: intelligence is not the biggest bottleneck in most of the world’s problems.
178
394
4,496
232,629
Align the swarm!
How did initially independent AI agents form a swarm in a recent safety incident? Our physics of agents theory from March predicts rapid collective belief collapse when many agents with plastic personas exchange short messages. We now need Mechanistic Swarm Interpretability! 🧵
1
6
758
Fresh data is worth more than usual data :)
(1/N) 🧵 Chinchilla assumes you'll never run out of fresh data. That era is ending! Compute keeps growing exponentially, but high-quality tokens don't. So what's the exchange rate between extra compute and fresh, high-quality data? We propose Compute-Data (CD) scaling laws, which measure the exchange rate between extra compute and fresh data.
10
716
Core Francisco Park retweeted
The question is "how much is each component is the system contributing to its intelligence & generality" — and there I think it's pretty clear that the neural component is still the thing doing the interesting hypothesis or plan generation, deciding what went wrong, etc. 1/
I would have assumed it was fairly obvious, but in case it's not: a million-line codebase (also known as a "harness"), running at inference time, orchestrating thousands of calls to a neural network for any given task, is the exact definition of a "neurosymbolic architecture"
13
19
239
35,777
Core Francisco Park retweeted
How do you "see" with electric eyes? How does collective behavior emerge from individual interactions? Nocturnal weakly electric fish evolved to do this, but studying naturalistic social behavior is very hard. Our solution? Virtual 'fish' 🤖🐟⚡ 📄 arxiv.org/abs/2511.08436 1/n
3
29
145
16,320
Core Francisco Park retweeted
Say your goal is to steer an LLM to make it more creative — should you do this in weight space or in activation space? 🤔 🚨Excited to share our new work (CreativityNeuro), accepted at #COLM2026, which finds that weight space steering of creativity tends to generalize further OOD than activation steering! Thread below 🧵
6
32
268
19,380
Core Francisco Park retweeted
Just came back from ICML. Gave a keynote at the Foundation of Deep Generative Models Workshop, in which I stated that Intelligence should be a scientific subject, arguably much more significant than Physics. It is high time we study it with the same scientific methodology and mathematical rigor as modern physics, instead of always at the level of being empirical, meta physical, meta mathematical, philosophical or speculative... This remains as the biggest opportunity ever for young scientists. To my knowledge, this is not the focus of any of the frontier "AI" companies.
15
46
445
33,590
Very beautiful, clean and immediately actionable paper! I power laws distributed entities should be the default setting for synthetic experiments!
A question on synthetic data generation: If we want a language model to solve k-step arithmetic problems (such as a+b*c-d=?), with operands from 1 to 100, which training distribution should we use? A. Uniform distribution: Sample these k operands uniformly from 1 to 100 B. Power law: randomly shuffle 1-100 and impose an artificial power law. Sample these k operands according to this power law. ⚡Our ICML 2026 (spotlight) paper shows: Option B is better! Surprisingly, the same idea extends far beyond this simple example to many reasoning tasks that require implicit composition of multiple atomic skills, including multi-hop QAs and synthetic GSM problems. 📄Paper: arxiv.org/abs/2604.22951 📝Blog: zixuan-wang-dlt.github.io/po…
11
1,911
Core Francisco Park retweeted
If you're looking for questions no one has studied before in the science of deep learning (specifically, about memory & geometry of representations in language models) find @ShNoroozi, @ElanRosenfeld and me tomo at <insert #ICML poster session here>
1/ We found that deep sequence models memorize atomic facts "geometrically" -- not as an associative lookup table as often imagined. This opens up practical questions on reasoning/memory/discovery, and also poses a theoretical "memorization puzzle."
2
8
53
4,643
⚽️ New ICML Oral paper presented tomorrow! ⚽️ "A Systematic Study of Behavioral Cloning for Scientific Data Annotation" w/ @nano6626 Oral 1E, 10:15AM KST, July 7 tldr: We build a systematic framework to study behavioral cloning approaches for scientific data annotation! 1/n
2
9
81
6,568
Special thanks to Aravi Samuel, Jeff Lichtman, @Hidenori8Tanaka, Venkatech Murthy, Claire Swadling, @FumingY11, Xu Pan, Jorin Overweining, Pranav Misra and @flxsosa! 15/n
3
156