IROS 2026 is ongoing. So here's the top 10 papers/posters you can try exploring..
1) Scaling Cross-Embodiment World Models for Dexterous Manipulation
Human hands + very different robot hands are represented as 3D particles, giving one world model a shared space for planning across embodiments.
arxiv.org/abs/2511.01177
2) AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models
A latent world model scores candidate robot actions during offline VLA post-training instead of requiring endless physical rollouts.
arxiv.org/abs/2603.08519
3) OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Models
Instead of forcing a VLA to learn every camera viewpoint, reconstruct the scene in 3D and render canonical views first.
og-vla.github.io/
4) RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation
A nice middle ground between cheap 2D video models and expensive full-3D world models: use a geometry-aware 2.5D representation.
arxiv.org/abs/2510.09036
5) TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
Touch tokens are activated when contact happens, instead of feeding tactile information into the VLA all the time.
arxiv.org/abs/2603.12665
6) AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint-Robust VLAs
Move the robot camera and the policy can break. AnyCamVLA re-renders the new observation toward the viewpoint the VLA was trained on.
arxiv.org/abs/2603.05868
7) LangGap: Diagnosing and Closing the Language Gap in VLAs
Some VLAs can score extremely well while barely using the instruction. LangGap keeps the scene fixed and changes the language to expose that shortcut.
arxiv.org/abs/2603.00592
8) VG3S: Visual Geometry Grounded Gaussian Splatting for Semantic Occupancy Prediction
Frozen vision foundation models provide geometric priors that help build better Gaussian-based 3D occupancy representations.
arxiv.org/abs/2603.06210
9) EquiBim: Learning Symmetry-Equivariant Policy for Bimanual Manipulation
If you mirror the scene and swap the robot’s arms, the action should mirror too.
Simple physical symmetry turned into a learning constraint.
arxiv.org/abs/2603.08541
10) Learning Tactile-Aware Quadrupedal Loco-Manipulation Policies
The robot does not just predict where its hand should move.
It also learns what the desired tactile interaction should feel like over time.
arxiv.org/abs/2604.27224