Interesting how embodiment and intelligence condition ond another
Robotics essentials: EMBODIMENT
A robot's intelligence is not only stored in its head. Part of it sits in the shape of its body.
The oldest proof is the passive dynamic walker: a pair of legs with no motors, no sensors and no controller at all, which walks down a shallow slope on gravity alone. The gait is not computed. It falls out of the geometry. Build the body right and it does some of the thinking for you, for free. The Cornell biped built on that principle walks at roughly human energy cost. Honda's ASIMO, computing every step, used about an order of magnitude more.
The flip side is the trap the field is in now.
A policy learns the body it grew up in: its joint count, its reach, its camera placement, its control rate. Move it to a different robot and it doesn't degrade, it dies.
Tested on robots they weren't trained on, every open vision-language-action model collapses to zero success. Not worse. Zero.
The escape plan is cross-embodiment: pool data from many robot bodies and train one model on all of them. Open X-Embodiment did it with 22 robot types and a million demonstrations. It works, unevenly, and how unevenly is a post of its own.
Which is the real argument for building human-shaped robots, and it is not about the hardware.
It is that the human body is the one morphology with a training corpus already sitting on the internet. Nvidia pretrained on 20,854 hours of human head-camera video; going from 1,000 to 20,000 hours more than doubled what the robot could finish. On the same task, one hour of human hand video beat one hour of robot data.