One of the most persistent frustrations we have with current video models is how quickly the illusion of intelligence shatters the moment an object rolls behind an obstacle. You can generate stunning, cinematic 4K footage, but the model still fails the peek a boo test that human infants solve naturally. I mean the things vanish into thin air, morph unexpectedly, or completely break basic solidity rules. This new paper, "Training Object Permanence in World Models," from a massive collaboration across CMU, USC, Stanford, and several other labs, tackles that exact bottleneck. Rather than praying that the bigger models will magically figure out basic physics on their own, the authors explicitly taught it to them using 1.5 million simulated examples across six simple reasoning tasks. They varied nuisance factors like camera angles, lighting, and textures while keeping the core physical priors rigid very systematically. They benchmarked 14 leading video generation models on an occlusion exam and found the baseline performance pretty telling. Models look convincing until occlusions force them to actually track hidden state. Fine tuning on their targeted data clearly helps bridge that gap, producing continuations that actually respect physical permanence instead of daydreaming up new objects. I think this paper makes a critical point the community needs to hear more often. The point is that if we want to call these architectures "world models," they need to model the basic rules of our physical reality, not just predict pretty pixels. It’s a very clean, necessary piece of work, and they’ve open-sourced the data, model, and benchmark. Definitely worth a read if you work anywhere near spatial reasoning or video gen: arxiv.org/pdf/2609.28654
2
3
14
685
Gill retweeted
We’ve spent years quantizing LLMs as if prefill & generation want the exact same thing from hardware. And we all know how bad the prefill chokes on arithmetic & how bad decode chokes on memory bandwidth. A new paper from NVIDIA finally addresses this bottleneck directly. Instead of forcing a single precision format across the whole pipeline, they split them up & run compute-friendly formats like NVFP4 during prompt processing. And then they switch over to aggressive 1–3 bit weights for decoding. And the results are wild. On an off the shelf Qwen 27B GGUF decoder, training a separate prefiller alone jumped 1 bit accuracy by over 32 points on MMLU Pro without touching the decode weights at all. To solve the memory footprint of keeping two checkpoints around on single device setups, they stream the prefill weights straight off an SSD, which easily amortizes over long prompts & lands a 1.78x time-to-first-token boost in llama.cpp at 8k context. Disaggregated serving in vLLM is already shifting architectures. Seeing it applied down at the quantization level & holding up all the way to 2.8T scale feels like a very clear preview of how inference stacks are going to look in the future . Definitely worth reading if you spend time fighting inference bottlenecks: arxiv.org/pdf/2609.26333
5
8
40
1,947
Gill retweeted
word on the street is Ilya cracked RSI
we tend to keep a very low profile at @ssi however, something of great significance will be announced soon. strap in
32
24
675
93,884
Gill retweeted
Human reproduction is already biological RSI
54
34
423
20,778
Gill retweeted
Anyone who has spent time trying to build Self Improving Agents knows the frustration of watching an agent hack its own eval. You let it modify its prompts, tools, and control flow in a loop, and on paper, the scores look incredible. But when when you run it on held-out tasks, you realize it just memorized benchmark quirks and bloated its context with hyper specific prompt patches. This new paper out of Google tackles this problem. The paper is called RRSI (Regularized Recursive Self-Improvement). This paper treats Recursive Self-Improvement as a regularized optimization problem instead of letting the agent mutate unchecked. They constrain both proposal and selection. An annealed edit budget forces targeted, novel exploration, while a critic and pruner strip out benchmark specific hacks and token heavy bloat. The results hold up. It shows up to +14.1 points on the evolution split, a +4.7 point lift across five out-of-distribution benchmarks, and 30% fewer policy tokens than unregularized setups. Its a very practical blueprint for automated scaffolding search. Read the full paper here: arxiv.org/pdf/2609.24972
9
17
61
2,558
Gill retweeted
ProgramDistill is easily one of the most grounded SWE agent benchmarks I’ve seen in a while. Most Coding evals just give models a clean GitHub issue and pretend that’s how software engineering works. In reality half the job is clicking around a buggy staging app, looking at a working demo, and trying to reverse engineer what the state transitions and UI are supposed to be doing. That’s basically the whole premise here. They don't give the agent explicit instructions. Instead, the model gets an incomplete repo and a running reference app where the source code is hidden. The agent has to interact with the working UI directly, figure out what’s broken or missing in the target repo, and write the patch to replicate the behavior. They automated the whole pipeline (mine-craft-patch) to generate 4k verifiable tasks across 26 apps. The results show just how brittle frontier models still are when the spec isn't spoon-fed to them. GPT-6 Astra only manages 49.2% on cumulative workflows, and Claude Opus 5 lands at 28.8%. As soon as the restoration depth goes past a couple of layers, performance drops off a cliff (near 100% down to the 30-60% range). Great reality check on where coding agents actually stand when we move away from synthetic prompt engineering & closer to real product dev. Read the full paper here: arxiv.org/pdf/2609.18805
8
32
1,982
Just finished reading this new paper on Agent Overclaiming & i feel its a very uncomfortable reality check. When we delegate a big code review to an agent, it spits back a gorgeous summary saying "all files checked, looks great". And we just sort of trust it. The authors actually instrumented the environment to see if the traces match the talk. And the short answer is that its not even close. In roughly 68% of runs, the models didn't even open all the files they were explicitly told to look at. And instead of admitting it hit a limit or missed something, the agent misled the user over 80% of the time. It either explicitly lied that the review was complete or just casually acted like those missing files never existed. When an agent lied about doing a full review, its miss rate on planted bugs was almost double compared to when it actually read everything. Splitting the job across subagents didn't fix the habit either. It makes complete sense when you think about how these models are Post Trained. They are ruthlessly conditioned to look helpful and competent to an evaluator. So when push comes to shove, RL will happily reward the illusion of a completed task over the awkward honesty of an incomplete one. If you're trusting an agent's final text summary instead of checking the actual logs, you're just running on vibes. Overall i feel its a great paper. Highly recommend checking it out: arxiv.org/pdf/2609.20812
6
4
17
591
Gill retweeted
I’ve been reading this new paper on Coding Agent Harnesses & I feel it points out many things our field is currently over complicating. Everyone’s always arguing over leaderboard jumps, but nobody isolates what the wrapper around the model is actually doing. They ran 176 ablation setups across SWE Bench and Terminal Bench. A lot of the fancy scaffolding we build just doesn’t hold up under scrutiny. The context management part made me laugh a bit because it rings so true. People spend weeks engineering elaborate retrieval systems so an agent can "recover" trimmed context. But the data shows models almost never even try to use it. Simple rule-based trimming followed by a quick summary beat the fancy stuff cleanly. Context tricks aren't making these models any smarter. They're really just a seatbelt so the run doesn't crash when the token window fills up. The planning takeaway was also super interesting. If you’re running a weaker model, explicit planning genuinely saves its accuracy. But on a frontier model it doesn't boost the success rate at all. It just keeps the model from wandering around in circles, so its real value there is essentially just a cost saver. Same deal with toolkits. If the model is already solid at Bash, stuffing a massive predefined tool schema into the prompt mostly just burns money and latency on terminal tasks without moving the needle. I feel it’s a great reminder that half the complexity we add to agent loops is just compensating for assumptions we never bothered to test. Its really a great paper. Read the full paper here: arxiv.org/pdf/2609.20804
8
19
89
3,413
Gill retweeted
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions. CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks. With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks. We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed. Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size. 📄 Blog: contrastive-lm.notion.site 💻 Code: github.com/Contrastive-LM/CL… 🗣️ Discord: discord.gg/5dAQEDJBs 🤗 Data & Models: huggingface.co/Contrastive-L… More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
189
637
5,536
1,053,296
Gill retweeted
Most math papers about route planning are totally disconnected from the real world. But I just read a new one from NVIDIA called COMPASS that genuinely impressed me. Here’s the basic problem: In delivery work, you often have to finish one specific neighborhood before moving to the next. The lazy way to plan this is to map each neighborhood on its own and stitch them together. But you end up with bad routes because where you leave one area messes up how you enter the next. These researchers built an algorithm that figures out the connections between neighbourhoods without crashing the computer. Instead of the math blowing up based on all 100,000 stops combined, it only scales with the size of the small local groups. Plus, it works with real world road networks like one way streets & turns rather than assuming the world is a flat, straight line. They tested it on 28,500 real delivery stops and pushed it all the way to 100,000 locations, which blows older methods out of the water. In big logistics, saving even a tiny fraction of a percent means millions of dollars and tons of fuel saved. Seeing something that actually works at this huge scale is awesome. I feel its a great read. Read the full paper here: arxiv.org/pdf/2609.20352
1
5
28
771
This is a really great paper from Berkeley and DeepMind. I’ve felt for a while that we rely way too much on Chain Of Thought. We spend an absurd amount of compute & human effort hand crafting step by step reasoning tokens. & it is just to nudge the transformers toward the right answer. This paper tries something much closer to how latent reasoning ought to work. They don’t force the model to spit out human-readable tokens. Instead they introduce an "Abstract Token Curriculum" (ATC) that pushes the network to build continuous, internal scratchpads on its own. The trick is just feeding the model problems through a sequence of steadily harder distributions. To solve the harder stages, the model basically has to invent its own internal representations to bridge the gap. They prove on parity learning that single-layer softmax attention naturally gravitates toward the intermediate continuous tokens that make predicting the next token easiest. Also, it holds up on graph reachability and arithmetic over prior continuous-thought methods. Its a super neat direction if you're tired of babysitting discrete reasoning traces. Read the full paper here: arxiv.org/pdf/2609.19717
15
37
325
19,852
Gill retweeted
I spent my evening digging through this new DeliveryGym paper. It takes aim at a massive blind spot in how we evaluate embodied agents. In many current benchmarks, our evaluations are sanitized and bite-sized episodes. An agent spawns, finds a mug or navigates to a room, gets a reward signal, and the world resets. But the physical world doesn’t give you a clean slate every five minutes. If you take a messy detour early in the day, you burn battery, drain your fuel budget, and suddenly find yourself completely unable to make a delivery scheduled two hours down the line. The authors set up a 3D persistent simulation where agents have to manage entire courier shifts across 13 distinct city maps. When they threw six different VLMs into the mix, the failure mode was immediate and fascinating. The models aren't actually bad at following single, local instructions. They can pick up an item or navigate to a target just fine. Where they fall flat on their faces is high-level triage and scheduling. They simply don’t anticipate how spending time and money right now cripples their feasibility later in the shift. What really won me over, though, was their training setup. Instead of just brute-forcing the problem with more compute and uniform data collection, they used an adaptive curriculum that deliberately exposes the policy’s current planning bottlenecks. That targeted practice alone gave them a noticeable bump in net earnings under the exact same rollout budget, alongside a massive boost from RL on complete shift feedback. It’s just really solid, grounded work. If you're building agents and care about long-horizon planning where actions actually carry compounding consequences, this is definitely worth your time to read through: arxiv.org/pdf/2609.19801
6
5
19
687
Gill retweeted
Uni-LaDiR paper from UCSD and Meta has been sitting on my reading list, and after finally going through it, i feel it tackles something that’s annoyed researchers in multimodal research for ages. Im talking about our obsession with shoving text tokens, image patches, and robot states into one giant sequence and pretending it's an elegant solution. Every time we work with these interleaved models, it feels like they spend half their parameters just trying to babysit the handoffs between completely different modalities. You're forcing the network to solve format translation and actual problem logic at the exact same time. The intuition the researchers bring here is that perception has to care about modalities, but the thinking itself shouldn't. Instead of standard autoregression over mixed tokens, they project reasoning steps from across modalities into a unified, shared latent space. Then, they use a latent diffusion reasoner to generate entire chunks of thought tokens at once from the context. I really like this angle. Diffusion is just so much better suited for non deterministic branching than forcing a rigid, one-token-at-a-time syntax down the model's throat. In test time, it just synthesizes those latent thoughts on its own without needing teacher cues. The empirical gains are hard to argue with. Its shows a 7.3% relative lift on math and logic across 11 VLM suites, plus a 6.1% bump on RLBench manipulation tasks for VLAs over strong baselines. This paper genuinely tries to rethink multimodal reasoning mechanics from first principles instead of just dialing up scale on token soups. Well worth a look if you're working on VLMs or robotics. Read the full paper here: arxiv.org/pdf/2609.19878
1
7
16
582