@theworldlabs shipped a world model you can aim. 🎯🌍
Most video models still treat camera movement as language. You describe a crane shot, cross your fingers, and pull the lever again when the drift ruins it.
🤳Atlas, out today from World Labs, takes camera geometry as a native input type. Position, angle, path. Up to one minute of video at 1440p that holds together for the whole run.
Under it: a multimodal autoregressive diffusion transformer, pretrained from scratch on text, images, video, and 3D.
An LLM builds its context out of tokens in a row. Atlas builds its context out of images pinned to positions in space. Every frame it has seen knows where it stands, so whatever comes next has to agree with the geometry and not just the vibe.
📸Drop two unrelated photos into that spatial context, place them apart in 3D, and the model invents the hallway between them.
➡️ Reconstruction from one to dozens of images, no capture rig, no hundreds of dense views
➡️ Two or three views already give faithful reconstructions, and on sparse-view benchmarks the generalist beats open-source models trained only to reconstruct
➡️ Depth maps are native, so worlds come out as point clouds or Gaussian splats, the same representation already running in Marble
➡️ Third-party human raters picked Atlas over recent video models on camera following in 75 to 94 percent of head to head votes, and the advantage grows as the trajectory gets more complex
Then the robotics layer:
Two large environments, captured on a cell phone, 24 frames each. Atlas reconstructs the space, then generates the RGB and depth a simulated robot's body-mounted cameras would see as it moves through.
The room and the robot's view of the room come from one model, so nothing gets handed off to a second system that can disagree with the reconstruction.
Manipulation goes further.
From a few casual recordings, Atlas helps build a scene where rigid, articulated, and deformable objects behave, then lets you swap the objects, the lighting, the positions, the motion.
Scanning spaces like these traditionally required elaborate and expensive equipment. Now it is a phone in a backpack.
Robots were never short on ambition. They were short on rooms to practice in.
🔗 Full technical post:
lnkd.in/gJ5UPaQu
🔗 Real-to-sim background:
lnkd.in/gZ_FZZwd
🔗 Early access:
lnkd.in/gz8YXUMw
🔗 Marble:
lnkd.in/g3KSKZ_p
Congratulations to
@drfeifei,
@jcjohnss ,
@chlassner ,
@BenMildenhall,
@KeunhongP,
@ychngji6,
@XRarchitect and the World Labs team, and to
@YunzhuLiYZ and Changxi Zheng on the robotics side.
⭐
@BotNewsAI botnews.ai for the latest in emerging tech