Building AI that learns by interacting with the world. Associate Professor @ MIT, leading the Scene Representation Group (scenerepresentations.org).

Cambridge, Massachusetts
I am open-sourcing DeckWerk, a multi-platform slide editor with stellar video support, real-time collaboration, and Agent support. I used to be on Linux but had to switch to Mac b/c there was no good enough slide editor - @OmarchyLinux & @dhh inspired me to vibe-code my own, so I'm finally back on Linux! Install from source via GitHub (installers coming soon!): github.com/vsitzmann/deckwer…
19
39
401
45,014
At Rhoda, we are constantly extending the commercial viability of robot automation - we have a long list of real customer tasks that robots can do today, with high autonomy. This is one of many :)
More industries. More real customer work. Introducing RhodaAtWork, a series on the customer workflows our robots are learning to handle across manufacturing and logistics. First up: unpacking, scanning and stacking component reels for electronics manufacturing.
1
4
75
8,279
Excited to speak at this workshop!!
The hottest debate of the year: do robots need 🌎 world models 🌎? Join us at CoRL 2026 for a scientific debate on the case for — and against — world models for robotics, featuring @chelseabfinn @vincesitzmann @YunzhuLiYZ @GeorgiaChal! 📄 Paper submissions are open now and close October 12. We welcome both research papers and position papers. 🏆 Submit your work for a chance to win an NVIDIA Jetson Thor or a Samsung Galaxy Tab S11 Ultra! Organized by @longhini_a @wenlong_huang Lasse Peters, @Jefferson_Aero , @PatkiSiddharth @jeff_ichnowski @drfeifei and myself. Thanks to sponsors @NVIDIARobotics and @Samsung!
1
41
5,714
I think DHH nailed it with many parts of this keynote - not just on the role of software engineering in general, but also as a healthy outlook on how AI might transpire in the future. I largely agree with his vibe here :)
It's pencils down, people. Writing code by hand is no longer an economically viable skill for most programmers at most companies. But the future of making software has never been brighter. Don't you dare black pill this beautiful moment! piped.video/vDjW_dRyKXY?si=6Fsf…
1
23
7,902
Accepted at NeuriPS - meet us in Sydney!!
Introducing MilliVid, our new method for long-context video generation! MilliVid creates videos that are consistent over long time spans, without using retrieval heuristics or 3D maps! (1/n) davidcharatan.com/millivid/#
1
7
136
13,083
This paper has been accepted at NeurIPS! Meet the team in Sydney!
We discovered that our latest Dataset Distillation project can be used to create some beautiful synthetic images based on an artist's body of work! Come see us at the @eccvconf Art Gallery starting today! Explanation and some of my favorites in thread below: 1/ (Claude Monet)
1
6
74
8,541
Introducing MilliVid, our new method for long-context video generation! MilliVid creates videos that are consistent over long time spans, without using retrieval heuristics or 3D maps! (1/n) davidcharatan.com/millivid/#
11
72
438
75,175
Now accepted to NeurIPS!!
1
378
I am open-sourcing DeckWerk, a multi-platform slide editor with stellar video support, real-time collaboration, and Agent support. I used to be on Linux but had to switch to Mac b/c there was no good enough slide editor - @OmarchyLinux & @dhh inspired me to vibe-code my own, so I'm finally back on Linux! Install from source via GitHub (installers coming soon!): github.com/vsitzmann/deckwer…
19
39
401
45,014
It's entirely vibe-coded with something like ~100 hours of my time put in (my weekend and night project) and many many more agent hours. I use it exclusively now instead of Keynote (which in the year of our lord 2026 still does not allow you to crop videos!!). There are definitely still bugs and rough edges, but I have used it in ~7 talks so far and it has not let me down, both remote and in-person - I think it's ready for prime-time :)
1
1
23
2,166
Oops the release build was 2 weeks old - if you installed it earlier, pls reinstall it to get the latest version 😅
2
1,110
Vincent Sitzmann retweeted
We recently came out with our blog post, Does Scaling Web-Video Pre-training Help Real Robots Do Real Work? This is a pretty exciting moment for robotics, because as far as I am aware, it is the first evidence that robot foundation models can scale on internet video data; not merely on teleoperation, UMI data, or egocentric data, but on the data you can truly find anywhere. It’s the necessary requirement for reaching massively powerful models.
At Rhoda, we care deeply about the science of pre-training for robotics. In one of the most rigorous studies of its kind, over thousands of trials and hundreds of hours of robot evaluations, we show how scaling web-video pre-training leads to better real-world robot performance.
1
4
54
6,674
Vincent Sitzmann retweeted
Stop by the #ECCV2026 Art Panel in Malmömässan C1 at 13:30 today (Saturday) to get a sneak peak at the tech behind these art distillations!
We discovered that our latest Dataset Distillation project can be used to create some beautiful synthetic images based on an artist's body of work! Come see us at the @eccvconf Art Gallery starting today! Explanation and some of my favorites in thread below: 1/ (Claude Monet)
1
4
18
2,128
Tongzhou has carefully evaluated the actual real-world scaling laws of Rhoda's web-scale video generative model - very cool study!
Does scaling pre-training on general web video improve a complex manipulation task in real deployment? We scale model size and pre-training compute, and test on one industrial task. Yes. The better a pre-trained model predicts web video, the better its post-trained policy. 🧵
1
31
4,712
Based on his pioneering work on dataset distillation, my student George has noticed that his most recent method (to be released soon, stay tuned!), besides being SOTA in dataset distillation, also creates stunning synthetic "composites" of an artist's body of work. Check it out! He was invited to display these at ECCV's art gallery and they came out amazingly :)
We discovered that our latest Dataset Distillation project can be used to create some beautiful synthetic images based on an artist's body of work! Come see us at the @eccvconf Art Gallery starting today! Explanation and some of my favorites in thread below: 1/ (Claude Monet)
1
8
130
11,291
I think this is my personal favorite :)
Replying to @GCazenavette
Thanks @elluba and @YSiglidis for organizing! You can find the gallery at the expo in Malmömässan near the t-shirt pickup. It's open to browse, and there will be guided tours at 11:00 and 17:00 on Thursday (today) and Friday and at 11:00 on Saturday. 4/ (Georgia O'Keefe)
4
1,219
I agree with Phil's take here: The progress of LLMs on controlling robots is quite interesting. Intuitively, this makes sense: controlling a robot is not so different from computer use, and an agent that is good at computer use is probably also good at controlling a robot and vice versa - credit to my student @RyuHyunwoooo to pointing that out to me originally!
Recently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots. I wrote a short blog post with my thoughts on the advent of these "robot-use agents." web.mit.edu/phillipi/www/wri… I think it's an important change in the trajectory of robotics!
8
22
192
42,956
One nuance is that it wouldn't surprise me if OpenAI had added a lot of robot and computer data to their training set, so it's not clear how much of this is "emergent" behavior!
2
1
22
3,295
On latency / speed concerns: I think LLMs and video models will become really fast over the last decade. It's good to remind ourselves that this technology has not been around for long! Even if the large models may not crack the hundreds of Herz that may be required for dexterous, dynamic control, I believe that smaller models that are actively learned via interaction with the real world just might.
8
1,426
At 10:20 am Malmö time, I will be speaking at the "X-Reason" workshop in the Palisades South room (turn right at registration, then down the stairs) about whether intermediate representations are important for embodied intelligence! #ECCV2026
2
37
4,188
This looks like a really cool product, congrats to the @theworldlabs team!
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
5
1
54
6,091