Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
Project:
research.nvidia.com/labs/dir…
Paper:
arxiv.org/pdf/2601.16163
Code:
github.com/nvlabs/cosmos-pol…
In this new work, researchers from Nvidia&Stanford developed a way to finetune Nvidia Cosmos video foundation model to serve as policy model, world model and RL value function model for robot manipulation tasks with no architectural changes.
- Inputs and outputs for Cosmos Policy: (1) inputs: task description text, current robot proprioception, current multi-view images (2) outputs: robot action chunk, future states (next robot proprioception, next images observations), RL value function (expected reward-to-go)
- Key idea to adapt video backbone without architectural changes: represent action chunk, future state and value function as latent diffusion frames so it naturally fits the video model’s latent diffusion process and also harnesses the model’s pretrained priors.
- Joint training in latent diffusion scheme: during each training step, sample (s, a, s’, V(s’)). (1) 50% batches for policy training: given s, predict a, s’, V(s’) (2) 25% batches for world model training: given s,a, predict s’, V(s’) (3) 25% batches for Value Function training: given s,a,s’, predict V(s’)
- Cosmos policy can be deployed as (1) a direct control policy (use a and discard s’, V(s’)) (2) a planning policy which uses s’ and V(s’) to do search with best-of-N sampling
- Authors show some good results both in simulation and real world bi-manual manipulation tasks.