Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects. Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
16
40
276
29,317
Are we experiencing a GPT moment in robotics? Our MolmoSpaces benchmark seems to suggest so. Check out @omarrayyann's thread to see the details 👇
GPT-Astra outperformed all open-source VLA/WAM baselines on a subset of our zero-shot MolmoSpaces-v1 benchmark! traces: orayyan.com/ms-results
3
4
42
18,221
GPT-Astra outperformed all open-source VLA/WAM baselines on a subset of our zero-shot MolmoSpaces-v1 benchmark! traces: orayyan.com/ms-results
4
16
184
27,553
It used fewer reasoning tokens than other VLM APIs (though the counts might not be directly comparable). I couldn’t run the full benchmark for these API models to include them on our leaderboard (molmospaces.allen.ai/leaderb…) because of the API cost :)
1
7
987
I think distilling the model’s spatial knowledge into a raw action space policy would be an interesting direction. its good vision understanding makes this verifiable enough to happen in an automated closed-loop way
8
829
Training for the robot olympics with MuJoCo’s new surfacevel for geoms
11
29
322
25,487
Omar Rayyan retweeted
A solid step toward humanoid loco-manip, super cool!
Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects. Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
1
3
32
5,874
Learning mobile manipulation is hard because collecting data for mobile manipulation is hard. But seems like @omarrayyann has found a promising way to crack this problem!
Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects. Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
4
23
3,359
Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects. Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
16
40
276
29,317
We follow this sim-to-real recipe to train two policies: a single-object Fetch policy that works across scenes, and an object-conditioned policy that takes the target object's name as an instruction. Both are two-system architectures with the decoupled SONIC controller acting as the low-level controller.
1
8
1,046
Check out our website for more videos and interactive simulation episodes: orayyan.com/fetchman. This project was done with advice from @YuchenCui1 and great collaborators @max_argus, Zhi Li, Chang Yu, Yuxin @chenfanfujiang. Paper: arxiv.org/abs/2608.17027 Code (soon): github.com/omarrayyann/Fetch…
1
12
1,067
We will be in the #RSS2026 poster session at 6:30 PM – stop by with all of your questions!
MolmoSpaces provides singular scale and diversity. We built a benchmark that puts that scale to use. MolmoSpaces-Bench evaluates zero-shot policies across thousands of environments previously unseen to them under systematic variation, providing insights that go beyond a success rate % More Below:
3
16
1,793
I’m at RSS in Sydney 🇦🇺 to present: - MolmoSpaces, Tuesday 6:30pm - Contact Anchored Policies, Wednesday 4:00pm - a (tbr) work on training general visual loco-manipulation policies in sim. If you work on anything from learning with off-domain data to sim-evals, let’s chat!
1
3
23
2,530
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data? Introducing Do as I Do: establishing a much needed correspondence between human videos and dexterous robot data. Some fun insights below: 🧵
11
65
365
96,851
Wall-OSS is now the #1 policy on the zero-shot MolmoSpaces evals. A lot of details in their paper, I recommend checking it out.
We are open-sourcing Wall-OSS-0.5. Pretrain Once, Act Anywhere. Wall-OSS-0.5 is a VLA model for real-world robotic manipulation, exploring whether pretraining alone can produce robot capabilities directly testable on physical hardware before task-specific fine-tuning. Key technical highlights: • Gradient-bridged co-training • Vision-Aligned RVQ Action Tokenizer • Action-Space Supervision • DMuon distributed optimizer In zero-shot real-robot evaluation, the pretrained checkpoint achieved task-progress scores above 80 on multiple tasks, including Block Sorting, Fruit Sorting, Ring Stacking, and Rope Tightening. Paper, code, blog, and uncut videos: x2robot.com/oss#resources
2
10
44
7,452
Omar Rayyan retweeted
Robotics is still data starved. Collecting high-quality robot demonstrations remains brutally slow and expensive. Introducing COBALT: A cloud-native teleoperation platform designed for large-scale robot learning. We are democratizing data collection by leveraging the hardware everyone already owns: the smartphone All you need is to download an app (today)! Read on for more!
29
51
399
104,849