One of the most persistent frustrations we have with current video models is how quickly the illusion of intelligence shatters the moment an object rolls behind an obstacle.
You can generate stunning, cinematic 4K footage, but the model still fails the peek a boo test that human infants solve naturally.
I mean the things vanish into thin air, morph unexpectedly, or completely break basic solidity rules.
This new paper, "Training Object Permanence in World Models," from a massive collaboration across CMU, USC, Stanford, and several other labs, tackles that exact bottleneck.
Rather than praying that the bigger models will magically figure out basic physics on their own, the authors explicitly taught it to them using 1.5 million simulated examples across six simple reasoning tasks.
They varied nuisance factors like camera angles, lighting, and textures while keeping the core physical priors rigid very systematically.
They benchmarked 14 leading video generation models on an occlusion exam and found the baseline performance pretty telling.
Models look convincing until occlusions force them to actually track hidden state.
Fine tuning on their targeted data clearly helps bridge that gap, producing continuations that actually respect physical permanence instead of daydreaming up new objects.
I think this paper makes a critical point the community needs to hear more often.
The point is that if we want to call these architectures "world models," they need to model the basic rules of our physical reality, not just predict pretty pixels.
It’s a very clean, necessary piece of work, and they’ve open-sourced the data, model, and benchmark.
Definitely worth a read if you work anywhere near spatial reasoning or video gen:
arxiv.org/pdf/2609.28654