I've been wanting to make 3D reconstructions not just realistic, but also **interactable** and **actionable** for years. Thanks to @XHongchi97338, we're now a step closer! Introducing DRAWER — a framework for the automatic construction of realistic, interactive digital twins.
Glad to introduce our #CVPR2025 paper "DRAWER", allowing one to create a realistic and interactable digital twin from a video of a static scene without any interactions with the environment. It unlocks many opportunities in gaming and robotics! Webpage: drawer-art.github.io
1
21
117
13,742
Wei-Chiu Ma retweeted
Ever suffered from making slides like I do? 😭 Aligning everything, fighting animations, trying to add cool interactive visualizations…while AI generated slides still look extremely confusing and overwhelming. So I vibecoded JostSlides ⭐ github.com/guangzhaohe/JostS… 🧵 (1/n)
8
8
45
3,660
Wei-Chiu Ma retweeted
Students sometimes ask me if it still makes sense, in this accelerating age, to pursue a PhD in AI. Perhaps counterintuitively, I think it's a great time to do so. I wrote up some thoughts on this here: web.mit.edu/phillipi/www/wri…
92
577
3,835
1,723,869
Wei-Chiu Ma retweeted
AI often decouples knowledge—which is about access to information—from understanding—which is about mastery. There is danger in settling for the former when what one needs is the latter. Understanding is harder to gain and quantify, but ultimately much more important.
1
7
346
Wei-Chiu Ma retweeted
Can 4D Foundation Models Remember? @AlexHe00880585, @ElorHadar, @weichiuma tl;dr: benchmark->visual memory in 4D foundation models arxiv.org/abs/2609.20819
1
10
42
2,815
Wei-Chiu Ma retweeted
We built our 3D coding harness earlier this year (arxiv.org/abs/2606.02580), back when none of the models we tested could really hit the bar. Revisiting it now with stronger models like Astra, we’re seeing a sharp threshold: below it, the harness helps a lot. above it, almost not at all. This is mostly the same across all model providers. A small warning for everyone building “agents”: all of today’s scaffolding will eventually get eaten by the base model. Also, we're not done playing with this. Sharing more interesting stuff soon :)
11
19
188
10,800
Wei-Chiu Ma retweeted
Getting the details of a sim2real stack is a huge pain, and attention to detail is key! We’ve been using Tao’s library at UW and would highly recommend :)
🚀 Looking for a reliable Franka controller with seamless sim-to-real transfer? Meet FrankaTwin: github.com/tsrobcvai/frankat… 🎮 Joint/Cartesian position & impedance + gripper 🎯 System ID: 3.6 mm EE / 28 mrad joint RMSE (sim vs. real)
5
79
9,226
Wei-Chiu Ma retweeted
I increasingly think future software will combine code for control flow with small neural programs for "fuzzy" judgments. That's what I've been exploring with ProgramAsWeights, which answers the question of where those neural programs come from: they are "compiled" from English descriptions. For example, I combined 30 neural programs with a decision tree to build a course website helper that looks like a chatbot but runs locally. Each small program handles a question like "Which specialist answerer should this be routed to", and code controls the overall flow. You can try building with it here: programasweights.com
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
13
21
178
13,881
Wei-Chiu Ma retweeted
Proxy Policy Steering (PPS) specializes VLA models to new tasks at inference time without accessing their weights. Instead of direct fine-tuning, it trains two lightweight proxy policies and uses their velocity-space residual to steer the frozen base's behavior.
2
10
64
3,970
Wei-Chiu Ma retweeted
How can generalist policies learn new skills without losing their pretrained prior? We introduce Proxy Policy Steering (PPS), which uses proxy policies to steer the frozen base policy at inference, adding task-specific behavior without touching its weights.
7
24
140
28,599
Wei-Chiu Ma retweeted
With all the Astra demos, there might be a feeling in robotics of a bitter lesson taking hold? I think this moment is rather the opposite. Astra works *because* it leverages traditional human-engineered methods as its tools. (...for now)
10
13
235
31,505
If you are at #ECCV2026, check out our #Wild3D workshop this afternoon! We have an amazing set of speakers: @vincesitzmann, @BenaimSagie, Georgios Pavlakos, @YuLunLiu790407, @RuohanGao1 You won't regret! 📌 13:30–17:00 @ Malmömässan B 🌎 3d-in-the-wild.github.io/ @eccvconf
Made with AI
5
15
901
Wei-Chiu Ma retweeted
Join us for the Functionality, Articulation & Interaction workshop at #ECCV2026 tomorrow! Our incredible lineup: - Andrea Vedaldi (@Oxford_VGG) - Wei-Chiu Ma (@weichiuma) - Evangelos Kalogerakis (@EvangelosKalog1) - Angela Dai (@angelaqdai) - Mikaela Angelina Uy (@mikacuy)
1
6
11
3,883
Wei-Chiu Ma retweeted
Here's my talk from the CVPR 2026 "Bitter Lessons" workshop earlier this summer. I've split it up into parts for the sake of discussion. Part 1: in which I wax poetic about my youth and force the audience to look at my dissertation results.
18
112
791
72,941
Wei-Chiu Ma retweeted
What seminal papers should every grad student in deep learning know? I'd like to update papers.baulab.info/ to add the most important new deep learning paper this year (or in the last few years). What should be added?
13
40
377
21,654
Wei-Chiu Ma retweeted
I’m incredibly excited to introduce @VeedaAI , which I co-founded with my incredible longtime collaborators @ZGojcic and @HuanLing6 . We started Veeda because we believe robotics will reshape the world—changing how we move people and goods, how we manufacture and build, and how we operate in the physical world. We also firmly believe that the scaling moment for Physical AI will come from robots learning through interaction with the world. Just like humans learn from interaction with their environments by trying, failing and trying again, until we succeed. And just like the capabilities of LLMs were truly unlocked when they started training in interactive environments and not only on vast amounts of internet text. For Physical AI, the central challenge is making this kind of interactive learning possible at scale. The real world is simply not a practical training ground for robots to learn through trial and error. It is unsafe, expensive, and time-consuming. To scale interactive learning, robots will need to learn in simulated reality. At Veeda, our sole mission is to build simulated reality for Physical AI. Our conviction is that this will become the critical infrastructure layer for all areas of robotics. As envisioned in pop culture, this means building “the Matrix” for Physical AI, where, through a virtual embodiment, robot intelligence can interact with an entirely virtual world in a million possible ways, learning from its own mistakes. We believe that World Models are the foundational technology that can make this possible. These generative models learn from massive amounts of sensor and physical-world data to reach the quality, diversity, and physical realism of the simulated reality that robots require. I could not be more excited to take on this challenge and build this transformative technology alongside an extraordinary team of engineers and researchers who have spent years working on some of the hardest problems across simulation, generative AI, 3D, robotics, and embodied intelligence. Together with our investors at @khoslaventures and @radicalvcfund , we intend to push the frontier of this technology and build the foundational infrastructure for interactive robotics learning and evaluation at scale. Come build this future with us. And to roboticists, help us understand what you need, we’re building this for you. veeda.ai
103
94
1,173
207,486
Wei-Chiu Ma retweeted
Was just as excited as everyone else to read this cool new paper, and it felt like a trip down memory lane! XM seems to be the same as IMLE, and the motivation and insights are also similar. Having worked on IMLE for years, I'll provide some context for my followers 👇. 1/n
We discovered a third pretraining axis beyond parameters and data: exploration. Scaling exploration monotonically improves existing models across images/video/language, and unlocks end-to-end generation. In the simplest case, it's just a for loop. Introducing Explorative Modeling. TLDR: - Gains from exploration grow with scale: 7%→36% as data scales, 13%→23% as parameters scale, and gains double at 3× the compute - Adding exploration to ~SOTA baselines improves data efficiency by 6.2×, FLOP efficiency by 4.1×, parameter efficiency by 47%, and hits a near-SOTA 1.43 unguided FID on ImageNet - Exploration lets you trade training compute for generalization, and scales how end-to-end your generative model is - End-to-end Explorative Models (XMs) match diffusion performance on control tasks with up to 256× less inference compute 🧵Thread:
24
94
1,000
324,775
Wei-Chiu Ma retweeted
I know this is supposed to be the dystopian future, but our ProgramAsWeights (programasweights.com) is basically this idea: compiling function descriptions directly into runnable neural programs (LoRA adapters), skipping code altogether. The part I find most exciting isn't replacing code, but implementing fuzzy functions that are easy to describe but hard to write using code, like repairing broken JSON, classifying if a message is urgent.
2
8
59
5,998
Wei-Chiu Ma retweeted
Play2Perfect learns contact-rich, precise assembly skills like screwing. The best part is watching the rollouts slowed down: the policy makes tiny corrections and recoveries that are essential for fast, dexterous manipulation. Check out Tyler’s thread for more details 👇
🤖 How can we teach dexterous robots to perform precise, contact-rich assembly? Introducing Play2Perfect: first learn to play with objects, then perfect the policy for tight insertion, multi-part assembly, and screwing. Sound on! 🔊 🧵👇
6
35
2,578
Wei-Chiu Ma retweeted
Teaching a robot shouldn't require humans to act like robots. Human demonstrations contain valuable signal for robot manipulation, but they aren’t directly transferable to robots. X-Diffusion learns from noisy human demonstrations while staying within the robot’s capabilities.
3
20
84
7,666
Wei-Chiu Ma retweeted
MeshFlow brings high-quality mesh with interactive speed. It's time to make mesh generation flow! Arxiv: arxiv.org/abs/2606.23489 Web: qiisun.github.io/MeshFlow/ Code: github.com/qiisun/MeshFlow
Excited to share MeshFlow — a new approach that can generate meshes with a fraction of seconds, while achieving state-of-the-art generation quality. Secret sources? Instead of autoregressive models, use equivariant flow-matching!
1
17
114
43,305