🍺 LagerNVS (CVPR 2026) 🍺 LagerNVS is a generalizable, feed-forward, real-time Novel View Synthesis network which - performs rendering in real time, - generalizes to in-the-wild data, - works with and without known source cameras, - sets a new state-of-the-art among deterministic methods, - can be paired with a diffusion decoder for generative extrapolation. LagerNVS shows that 3D biases are useful for Novel View Synthesis but explicit 3D representations are not required to achieve them. We use 3D biases in (1) architecture design and (2) pre-training: (1) In NVS with explicit 3D representations (3DGS, NeRF) reconstruction is typically difficult and slow, but rendering is much faster and simpler. We mimic this process in the network design: we use a large (1B params) encoder and a small, lightweight decoder (ViT-B). This allows increasing the network capacity while still achieving real-time rendering. (2) The encoder, initialized from VGGT, was pre-trained with 3D reconstruction objectives, making the initial features 3D aware. Both substantially improve performance. Project page: szymanowiczs.github.io/lager… Code: github.com/facebookresearch/… Paper: arxiv.org/abs/2603.20176 Models: huggingface.co/collections/f… Work done with @jianyuan_wang @MinghaoChen23 Christian Rupprecht and Andrea Vedaldi
7
32
208
33,017
Two personal updates! > I defended my PhD from @Oxford_VGG this week 🎉 Many thanks to my examiners @vincesitzmann and Andrew Zisserman for insightful questions and a fantastic discussion, and to my supervisors Andrea Vedaldi and Christian Rupprecht for an amazing 4-year ride together!! > Two months ago I moved to NYC to join @holynski_'s team at @GoogleDeepMind, working alongside @Jimantha and an absolutely stacked team. Excited for this new chapter!
30
1
242
11,218
Stan Szymanowicz retweeted
Poland has acquired a stake in artificial intelligence voice company ElevenLabs, as the country hopes to become a major European AI hub bloomberg.com/news/articles/…
33
93
602
1,065,026
Stan Szymanowicz retweeted
Introducing ABC: open data, training, and infrastructure for robotics. We release the largest teleop dataset to date, and extensively investigate design decisions, pretraining, and post-training techniques. @arthurallshire @Cinnabar233 @adamrasb @redstone_hong @davidrmcall
34
113
735
367,621
Timely
Just in time: the Best Paper Award in Robot Learning #ICRA2026 uses 3D camera pose to improve policy learning. Pretty straightforward: robots live in a 3D world. Image credit: @CSProfKGD
1
16
2,150
Stan Szymanowicz retweeted
Just in time: the Best Paper Award in Robot Learning #ICRA2026 uses 3D camera pose to improve policy learning. Pretty straightforward: robots live in a 3D world. Image credit: @CSProfKGD
Have to disagree that robotics should not care about 3D. Robots operate in a 3D world. Hand-eye calibration, 3D perception, grasp planning, motion planning, and contact-rich manipulation all rely on 3D geometry. 3D remains fundamental to robotics.
3
26
197
33,552
Heading to @CVPR today, excited to see friends old and new! Ping me or come say hi if you want to chat about physical AI I’ll also be presenting LagerNVS on Saturday morning, might even risk a live demo
15
547
Stan Szymanowicz retweeted
📢📢📢 Velox 🚀: Learning Representations of 4D Geometry and Appearance In our #CVPR2026 paper, we introduce a method for learning a native 4D representation, useful for many downstream tasks, such as video-to-4D, 3D tracking, cloth simulation, and others! 🌐: apple.github.io/ml-velox 📝: arxiv.org/abs/2605.04527
7
54
179
22,433
Stan Szymanowicz retweeted
Gemini Omni lets you step into any world you can imagine, bringing text, images, audio, and video directly to life.
Made with AI
4
1
33
1,628
Stan Szymanowicz retweeted
Excited to share what we are building -- Genie experience grounded in real-world street view. Try it out at labs.google/fx/projectgenie
Project Genie is a @GoogleLabs experiment that lets you simulate dynamic worlds you can navigate in real time with Genie, our general-purpose world model. Today, we’re connecting Project Genie to nearly 20 years of Street View data from Google Maps — so you can now build interactive spaces based on real-world locations. Street View imagery in Project Genie is available now for places in the U.S., and will expand to more locales over time. #GoogleIO
Made with AI
8
42
424
67,999
Stan Szymanowicz retweeted
My whole life, I've wanted to be an elephant riding a motorcycle through my hometown. Now, it's finally possible.
Real-world models are here! Stoked to share how we're bringing real-world locations to life by integrating Street View into Genie. Try it now at labs.google/fx/projectgenie and read the blog for more info: blog.google/innovation-and-a…
13
21
437
76,496
Stan Szymanowicz retweeted
Real-world models are here! Stoked to share how we're bringing real-world locations to life by integrating Street View into Genie. Try it now at labs.google/fx/projectgenie and read the blog for more info: blog.google/innovation-and-a…
18
94
626
229,443
Stan Szymanowicz retweeted
Introducing VGGT-Ω: scaling feed-forward reconstruction across static and dynamic scenes, and studying whether the learned geometric representations transfer beyond reconstruction.
14
102
620
783,723
We made an interactive client-server viewer for LagerNVS with @JonathonLuiten! You can now interactively explore scenes from just a photo capture - no optimization, no 3D Gaussians, just load your images, run the model on a cloud GPU and stream the renders to your local browser. Check out the video below for some spaces I recently captured in Oxford, London and beyond!
5
26
175
17,329
Stan Szymanowicz retweeted
Feed-forward 3D reconstruction should not be limited to predicting one Gaussian per pixel. We introduce TokenGS, which uses learnable tokens to decouple the 3D Gaussian prediction from the image resolution and the number of input views. #CVPR2026Highlight [1/6]
6
44
246
45,782
Stan Szymanowicz retweeted
I’ve decided to leave OpenAI. Below is the note I shared with my team. Building Sora zero-to-one with you all has been the honor and adventure of a lifetime. As this team knows well, one of the best parts of working on video is that you can see the scaling behavior with your own eyes, and there have been a lot of oh-shit moments over the years. About a month into Sora, way back when this was just a tiny two-man effort, we saw a sample where a land shark swam by a bunch of intricate cacti in a desert (it was a weird prompt), and the details of each cactus were perfectly preserved once the shark passed. We had never seen object permanence like this in any video model. That’s when we knew we were onto something. It’s amazing how much the Overton window has shifted on video models. OpenAI has a high tolerance for crazy moonshots, but even back in July 2023 there was a ton of skepticism that high-fidelity 1080p multi-shot generation was achievable within a year given the state of video in the broader industry. We managed to get there 7 months later. While the original Sora ignited a huge amount of investment in video across the industry, it took the next generation of models with Sora 2 for the broader public to understand the transformation happening. I’m proud of all the sleepless nights before and after the launch this team endured in order to deploy the technology in a responsible way and help steer societal norms. I am immensely grateful to Sam, Mark, Aditya and Jakub for fostering a research environment that allowed us to pursue ideas off-the-beaten path from the company’s mainline roadmap. It’s tempting in life to mode collapse to the most important thing, but cultivating entropy is the only way for a research lab to thrive long-term, and Sam deeply understands this. Sora was a project that could not have happened anywhere but OpenAI, and I will always deeply love this place for that. I’m going to miss this team a lot, but there are great things on the horizon for all of you. I will always be in your guys’ corner cheering you on. Love, Bill
213
198
3,834
536,500
LagerNVS is a highlight! Catch us in Denver!
🍺 LagerNVS (CVPR 2026) 🍺 LagerNVS is a generalizable, feed-forward, real-time Novel View Synthesis network which - performs rendering in real time, - generalizes to in-the-wild data, - works with and without known source cameras, - sets a new state-of-the-art among deterministic methods, - can be paired with a diffusion decoder for generative extrapolation. LagerNVS shows that 3D biases are useful for Novel View Synthesis but explicit 3D representations are not required to achieve them. We use 3D biases in (1) architecture design and (2) pre-training: (1) In NVS with explicit 3D representations (3DGS, NeRF) reconstruction is typically difficult and slow, but rendering is much faster and simpler. We mimic this process in the network design: we use a large (1B params) encoder and a small, lightweight decoder (ViT-B). This allows increasing the network capacity while still achieving real-time rendering. (2) The encoder, initialized from VGGT, was pre-trained with 3D reconstruction objectives, making the initial features 3D aware. Both substantially improve performance. Project page: szymanowiczs.github.io/lager… Code: github.com/facebookresearch/… Paper: arxiv.org/abs/2603.20176 Models: huggingface.co/collections/f… Work done with @jianyuan_wang @MinghaoChen23 Christian Rupprecht and Andrea Vedaldi
1
6
37
2,668
Stan Szymanowicz retweeted
Check out our recent work on generalizable real-time novel view synthesis. 🥳🥳 It suggests that explicit 3D representations may not be necessary, as long as the right structural biases are learned.
🍺 LagerNVS (CVPR 2026) 🍺 LagerNVS is a generalizable, feed-forward, real-time Novel View Synthesis network which - performs rendering in real time, - generalizes to in-the-wild data, - works with and without known source cameras, - sets a new state-of-the-art among deterministic methods, - can be paired with a diffusion decoder for generative extrapolation. LagerNVS shows that 3D biases are useful for Novel View Synthesis but explicit 3D representations are not required to achieve them. We use 3D biases in (1) architecture design and (2) pre-training: (1) In NVS with explicit 3D representations (3DGS, NeRF) reconstruction is typically difficult and slow, but rendering is much faster and simpler. We mimic this process in the network design: we use a large (1B params) encoder and a small, lightweight decoder (ViT-B). This allows increasing the network capacity while still achieving real-time rendering. (2) The encoder, initialized from VGGT, was pre-trained with 3D reconstruction objectives, making the initial features 3D aware. Both substantially improve performance. Project page: szymanowiczs.github.io/lager… Code: github.com/facebookresearch/… Paper: arxiv.org/abs/2603.20176 Models: huggingface.co/collections/f… Work done with @jianyuan_wang @MinghaoChen23 Christian Rupprecht and Andrea Vedaldi
3
18
3,116