Building AI that learns by interacting with the world. Associate Professor @ MIT, leading the Scene Representation Group (scenerepresentations.org).

Cambridge, Massachusetts
I am open-sourcing DeckWerk, a multi-platform slide editor with stellar video support, real-time collaboration, and Agent support. I used to be on Linux but had to switch to Mac b/c there was no good enough slide editor - @OmarchyLinux & @dhh inspired me to vibe-code my own, so I'm finally back on Linux! Install from source via GitHub (installers coming soon!): github.com/vsitzmann/deckwer…
19
39
401
45,019
I am thrilled to join rhoda.ai @RhodaAI as an advisor, where I am helping harness the abilities of large-scale pre-training and video models for robotics, putting many of my lab's research learnings of the past few years into practice! I will be in-person at the Mountain View office for part of the summer - reach out if you want to chat :)
21
17
486
56,381
Quantitatively, this shows up as a dramatic improvement of PSNR and FVD over rollouts of hundreds of frames. Note that even a diffusion model with perfect memory can't achieve a flat consistency curve, as the model has to generate unseen scene content! (5/n)
1
2
13
1,507
We are very excited about the results: we show that on Minecraft, our model can memorize 3D scene geometry for *hundreds* of frames, without retrieval or expert-crafted 3D map heuristics, at the same token budget as a conventional five-frame-context diffusion model! (4/n)
1
2
16
1,738
Unlike existing variable-compression methods (e.g., FramePack), we not only variably compress the past, but also generate the future from coarse to fine. This is what rollout with our model looks like: we first denoise a large chunk of the most compressed latents, then superresolve them with a lossless recent context and progressively more compressed past frames. Every rollout step uses the same diffusion transformer weights! (3/n)
1
5
44
8,434
The core idea is to use a hierarchical tokenizer that encodes frames into varying numbers of tokens for variable compression rates. We then build a coarse-to-fine diffusion model that only keeps the most recent frames at full token budget and uses compressed past frames! (2/n)
1
3
44
5,145
Introducing MilliVid, our new method for long-context video generation! MilliVid creates videos that are consistent over long time spans, without using retrieval heuristics or 3D maps! (1/n) davidcharatan.com/millivid/#
11
72
438
75,178
This was @BoyuanChen0's final PhD effort, and he wanted to come up with a set of tasks that would be *truly* test generalization, so him & the team asked friends & family to record themselves executing random manipulation tasks! The results are really encouraging :) (4/n)
1
10
1,132
While this idea is of course not new, making this work with video models is difficult! Diffusion Forcing and History Guidance are key parts, allowing us flexible prompting of the video model plus the required prompt adherence and video consistency. (3/n)
1
9
1,269
This application was our original motivation for diffusion forcing - the idea is that task planning is akin to using a *learned* internal world model to imagine what the completion of the task would entail, then implement that plan. (2/n)
1
9
1,250
Introducing the Large Video Planner, a video generative model that serves a powerful robotic planner that strongly generalizes to even unseen (!) tasks, outperforming both SOTA VLAs and prior video gen planners! (1/n)
3
13
163
14,025
Despite never seeing ground truth poses during training and not making use of ANY 3D inductive bias (no SE(3), no Plucker, no splats – just transformers!), probing experiments show XFactor’s pose latent correspond to the ground truth 3D camera poses  (9/n)
1
5
412
We train XFactor on diverse large-scale, real-world datasets at both the scene and object level. Our method significantly outperforms prior self-supervised models RUST & RayZer in terms of transferability by large margins while achieving high quality NVS.  (7/n)
1
4
465
In addition, we propose a novel self-supervised NVS training objective which explicitly promotes transferability, and introduce a representation learning-inspired augmentation strategy for training with real-world video. (6/n)
1
3
442
We identify the concept of TRANSFERABILITY: the same pose representations lead to the same camera poses across scenes, as the key criterion for true NVS. We introduce True Pose Similarity (TPS) to measure it - a metric that quantifies how well poses transfer between scenes (4/n)
1
4
528
In contrast, XFactor is the first self-supervised and geometry-free method which admits fully transferable latent poses: a given set of pose latents will render the same camera trajectory in any scene! (3/n)
1
4
607
We find that while prior pose-free NVS models like RayZer predict SE(3) camera poses, these poses are entangled with scene content – the same “poses” create different camera trajectories in different scenes. This is not true NVS, as it does not afford camera control! (2/n)
1
6
819
Introducing XFactor: the first pose- and geometry-free method capable of true Novel View Synthesis (NVS). We re-think NVS and the concept of camera poses completely without concepts from multi-view geometry as a pure representation learning problem! mitchel.computer/xfactor/ (1/n)
2
24
150
9,144
Besides his academic accomplishments Boyuan contributed his very own "Boba Shop" at MIT. That makes it doubly sad for him to leave, but he *did* organize a next generation of Boba Tea generators!
1
10
1,238
Finally get to tweeting about this: @BoyuanChen0, my student co-advised with @RussTedrake, graduated recently! Boyuan has done groundbreaking work on video generative modeling and video as the "language" of robotics. He is off to OpenAI where I am sure he will do amazing things!
5
3
123
19,031
Really cool! Loved the below paragraph in particular, a cool way to think about gradients!
I ran this experiment to show that duality-based optimizers like Muon are not only *fast* but also *numerically different* to vanilla gradient descent. In particular, the weights move a qualitatively different amount in the same number of training steps. (1/4)
4
9
90
14,526
The learned Jacobian Field is interpretable: we can visualize it by color-coding sensitivity to the different control channels. In this way, we can see that the Jacobian field correctly discovers the 3D kinematics of the robot. 9/n
1
5
1,148
Neural Jacobian Fields are trained completely self-supervised, from multi-view videos of a robot executing random commands - no human labels or intervention. At test time, a *single* image suffices to reconstruct them, for closed-loop control from a single RGB camera. 8/n
1
5
1,249
The Jacobian Field can be directly used for inverse dynamics control. Given desired motions, our model solves for the corresponding control command at interactive speeds. 7/n
1
4
1,162
Neural Jacobian Fields enable us to equip any robot with vision-based control, irrespective of its sensors, fabrication, material, or actuation. Each point is mapped to its “system Jacobian”, which maps a change in motor commands to the 3D motion of that point! 6/n
1
6
1,232
Our approach enables this pneumatic bio-inspired hand to perform physical tasks and allows the $220 poppy arm to draw “MIT” in the air, all with using a single RGB camera as the only sensor! None of these motion trajectories are prescribed as part of our training data. 5/n
1
6
1,287
Conventional robots are “blind” when it comes to state estimation! This constrains their design to be “simple” such that an expert can design a model & requires expensive manufacturing & sensors. But bio-inspired, multi-material, and cheap robots cannot be modeled easily! 4/n
1
7
1,477
Introducing Neural Jacobian Fields, robot 3D kinematic models learned only from vision! They can model & control robots from just a single RGB camera, even those w/ intractable kinematics & no embedded sensors such as soft, 3D-printed pneumatic hands! sizhe-li.github.io/publicati… 1/n
2
78
499
54,290
The idea is simple: You take any sequence NN (transformer, RNN, …) and train it to diffuse sequences. Crucially, however, you add *independent, random* noise levels to each token! Intuitively, this randomly masks tokens to varying degrees! (2/n)
1
1
14
1,818
A simple example: confronted with pairs of shifted images, NIso approximately recovers the Laplace operator with its shift-equivariant sine-cosine eigenfunctions. In this basis, our functional map is approximately block-diagonal, preserving spatial frequencies under shifts! 6/n
1
1
402
Specifically, we ask that these transformations manifest as functional maps – linear transformations between functions. To enforce computationally tractable structure, we require that our functional maps are isometries with respect to a learned operator in the latent space. 4/n
1
3
436
Real-world geometry and 3D vision tasks are full of challenging symmetries that defy analytic expression. Camera motion, for instance, which is a “nice” SE(3) transform in 3D, becomes dense optical flow in image space w/o tractable group structure! 2/n
1
1
697
We can use FlowMap itself to supervise & pre-train the depth estimator! Pre-training leads to better results & faster convergence, but not strictly necessary—it works even without any pre-training! The key is “patch-match” regularization: similar RGB patch → similar depth. 12/n
1
8
2,066
However, solving for depth, poses and intrinsics as free variables via gradient descent does not work well (see the paper for why!). Instead, we reparameterize both poses and intrinsics in terms of depth and optical flow, *leaving only depth as a free variable*! 10/n
1
6
1,247
FlowMap minimizes a “camera-induced correspondence loss.” When a camera moves through a static scene, that motion induces correspondences on the image sensor according to the scene’s geometry, the camera motion, and the camera intrinsics, which we supervise with point tracks 9/n
1
7
1,322
Here are some point clouds reconstructed from FlowMap on popular scenes - it really works very robustly!! 5/n
1
2
7
1,747
FlowMap is a major step towards solving that problem: it is a fully differentiable, self-supervised structure-from-motion method! From only off-the-shelf point tracks / optical flow, FlowMap performs SfM that outperforms Colmap’s on Gaussian Splatting Novel View Synthesis! 4/n
2
3
15
2,684
Introducing “FlowMap”, the first self-supervised, differentiable structure-from-motion method that is competitive with conventional SfM like Colmap! cameronosmith.github.io/flow… IMO this solves a major missing piece for internet-scale training of 3D Deep Learning methods. 1/n
11
101
606
128,708
This project was really fun, just following a quirky idea - an example for where great students can take you :) I loved the idea of shape-shifting robots so much that we had to do it after Boyuan brought it up in a conversation! 9/n
2
2
7
1,777
Lastly, we release “DittoGym”, a comprehensive benchmark with eight diverse tasks that can test shape-shifting robots, ranging from robot shape matching, to locomotion and manipulation. 8/n
1
2
461
We tackle this with a fully-convolutional policy framework and a coarse-to-fine curriculum. Our policy ingests images of the robot and outputs a dense action grid. This coarse-to-fine curriculum is key to making this high-dimensional control problem tractable. 7/n
1
1
11
2,078
We then show how we can learn a control policy capable of individually managing each particle in shape-shifting robots w/ reinforcement learning. This is challenging, as the action space is very high-dimensional, but necessary, as we require fine-grained shape control! 6/n
1
4
1,045
We model different plastic-elastic materials via Cauchy Stress and the von Mises yield stress: If strain exceeds the yield stress, the robot deforms plastically, thereby shape-shifting. Below the threshold, deformation is elastic, preserving the robot’s morphology! 5/n
1
2
1,020
We model shape-shifting soft robots via the material point method, which combines particle- and grid-based physics. We parameterize actions as a “muscle field” that generates forces on the grid - these forces propagate to robots' material points, thus actuating the robot! 4/n
1
1
6
1,049
We were inspired by the recently demonstrated “magnetic slime robot”, a chunk of magnetic goo that can be actuated by magnetic fields to shape-shift to perform tasks! piped.video/6jmt7uTiXVA?feature… 3/n
1
1
4
1,394
Excited to introduce DittoGym @ ICLR, in which we study the control of a neat new kind of robot: soft shape-shifters! This is work done by @SuningHuan44558 during his visit at my group at MIT, jointly with my student @BoyuanChen0! Project page: dittogym.github.io/ 1/n
3
26
144
30,020
I also want to highlight related concurrent work by friends at OxfordVGG: szymanowiczs.github.io/splat… - looks really cooll! @StanSzymanowicz, @chrirupp and Andrea Vedaldi - check it out! 10/n
1
9
1,269
Further, our 3D reconstructions are directly exportable as 3D Gaussians, making them editable and interpretable! 9/n
1
6
595
All in all, our method significantly outperforms all baselines on scene-scale, two-view 3D reconstruction, while being orders of magnitude faster and requiring less memory! 8/n
1
6
624
Next, we dealt with local minima that famously arise in primitive-based representations. Instead of directly predicting the positions of pixel-aligned Gaussians, we instead predict their probability density along a ray! 5/n
1
7
701
We build a multi-view epipolar line transformer that can infer scale given two input images, which we show easily addresses this challenge! 4/n
2
9
781
We investigate two core questions in this paper. First, the question of varying scale: for all scenes reconstructed via COLMAP/SfM, the scale of poses varies across scenes. That means that single-image reconstruction cannot succeed, as the neural net cannot guess the scale! 3/n
1
1
10
951
Introducing pixelSplat: feed-forward Gaussian splats from image pairs! Led by @DavidCharatan and @sizhe_lester_li, collaborating with @taiyasaki! We propose a memory-efficient, fast and editable alternative to pixelNeRF based on 3D Gaussian Splatting! davidcharatan.com/pixelsplat… 1/n
8
45
269
47,675
How can we learn to generate 3D scenes directly with diffusion models if we only have images, no ground-truth 3d scenes? Ayush, Tianwei and George will tell you at our poster “diffusion with Forward Models”, #202!
16
143
13,651
Find us at poster 226 where Cameron is presenting his cool work “FlowCam”!
1
4
81
10,238
📢📢📢 Code for our 3D generative model that learns to generate 3D appearance and geometry from just a single image is out now! It's trained just from real-world multi-view images, and generates scenes directly w/o score distillation! github.com/ayushtewari/DFM/
1
17
133
16,990
The samples are *truly* diverse. Note that each sample here is a full radiance field, from which you could - at any point - extract the pointcloud. And they vary widely in the unobserved regions! 9/n
1
6
1,029
This works on *real-world* scenes in RealEstate10k and Co3D, and significantly outperforms score-distillation based approaches! This is the first time that any 3D generative model trained with images can sample from the distribution of such complex 3D scenes! 8/n
1
7
1,157
This is difficult, b/c we never observe ground-truth 3d scenes - we only observe 2D images! We propose a new diffusion model that can nevertheless learn to directly generate 3D scenes, by integrating the differentiable renderer into each denoising step. 6/n
1
9
1,230
Conventional, non-probabilistic models such as pixelNeRF that reconstruct a 3D scene from a single image generate blurry results for any parts of the scene that were not observed in the input image. 3/n
1
7
1,697
Introducing “Diffusion with Forward Models”, 𝗮 𝗺𝗼𝗱𝗲𝗹 𝘁𝗵𝗮𝘁 𝗰𝗮𝗻 𝗴𝗲𝗻𝗲𝗿𝗮𝘁𝗲 𝗱𝗶𝘃𝗲𝗿𝘀𝗲, 𝗿𝗲𝗮𝗹 𝟯𝗗 𝘀𝗰𝗲𝗻𝗲𝘀 𝗳𝗿𝗼𝗺 𝗮 𝘀𝗶𝗻𝗴𝗹𝗲 𝗶𝗺𝗮𝗴𝗲, 𝘁𝗿𝗮𝗶𝗻𝗲𝗱 𝘄𝗶𝘁𝗵 𝗶𝗺𝗮𝗴𝗲𝘀 𝘄/𝗼 𝗮𝗻𝘆 𝟯𝗗 𝗱𝗮𝘁𝗮! diffusion-with-forward-model… 1/n
14
89
474
88,803
Code and trained models now available here: github.com/cameronosmith/Flo… Can't wait to see what folks will do with this - please reach out, we'd love to hear about it :)
2
12
2,159
Our model can be fine-tuned on OOD sequences with our simple re-rendering losses. Here we apply our RealEstate10K model to a Tanks and Temples scene, plotting estimated poses before and after fine-tuning. 12/n
1
7
956
Given this 3D motion field, we solve for an SE(3) pose via a differentiable, weighted least squares solver. Weights predicted by our model and allow us to be robust to faulty correspondences, e.g. from dynamic objects, specularities, occlusions, and incorrect optical flow. 10/n
1
6
1,239
We outperform prior work on pose-free novel view synthesis: Check out comparisons with a baseline that regresses camera poses via a CNN (Video Autoencoder), and a recent approach which uses latent poses (RUST), on the task of video reconstruction. 7/n
1
8
1,123
In FlowCam, we work towards lifting the pose bottleneck by having our model estimate both 3D geometry and camera poses. Our key innovation in robustly estimating poses is to first lift optical flow into 3D scene flow via our generalizable 3D renderer, then solve for pose. 5/n
1
14
1,420
This work was led by @omcamsmith, collaborating with @du_yilun and @_atewari, at my Scene Representation Group @MIT_CSAIL. I can’t overstate how impressed I am by Cameron’s work on this one, he really took it above and beyond! More cool stuff coming: scenerepresentations.org/ 2/n
1
2
17
2,146
Introducing “FlowCam: Training Generalizable 3D Radiance Fields w/o Camera Poses via Pixel-Aligned Scene Flow”! We train a generalizable 3D scene representation self-supervised on datasets of raw videos, without any pre-computed camera poses or SFM! cameronosmith.github.io/flow… 1/n
2
89
464
88,624
Second, a lightweight cross-attention renderer. Our renderer doesn’t sample points in 3D to then project them onto context images. Instead, we *linearly* sample features on the ray’s epipolar lines, enabling faster & cheaper rendering with fewer samples.
1
1
12
1,951
We found two critical improvements. First, inherent ambiguities in monocular depth lead to blurring in methods that encode context images separately. We leverage a multiview ViT image encoder that helps resolve depth ambiguities and alleviates blur.
2
7
1,642
Methods such as PixelNeRF can synthesize novel views given few input images. However, they are limited to simple scenes and small baselines. In our CVPR paper, we present a method for high-quality novel view synthesis given only two distant observations: yilundu.github.io/wide_basel…
1
33
205
31,044