The model is not the part that’s stuck
Everyone’s shipping bigger world models,better planners, nicer humanoid demos & the missing piece is usually a fresh ground-level view of a real place Language models had the open web,robots don’t
That’s the bet
@vangrid_io is making
Phone in someone’s pocket, bounty for a specific spot,clip gets fingerprinted & anchored on Base so a buyer can check the origin. Not a new architecture slide,a way to order the room the model has never seen
@drfeifei has said the robotics data gap versus language is the obvious one
Huang called this the ChatGPT moment for physical AI. Both still leave the same problem: the stairwell,the dock, the warehouse aisle isn’t sitting in a scrape
Street View is late,a fleet car is expensive,a random upload has no receipt,on-demand capture with an onchain fingerprint is the boring fix
If the next demo fails,I’d look at the training set before the weights
Where do you think the real bottleneck is right now, the model or the room it’s never filmed?