mimic is the frontier lab for robot dexterity

Zurich / SF
Pinned Tweet
Introducing FLUX-mimic, a next-generation Video-Action Model for general purpose dexterity, developed in partnership with @bfl_ai. Late last year we published mimic-video and introduced Video-Action Models (VAM): a new family of robotics foundation models built on top of video generation models. We showed that robot control reduces to visual prediction, and that robot capability is downstream of improvements in video modeling accuracy. The obvious implication was that advances in the video modeling frontier would directly translate to increased capabilities in end-to-end robot learning. FLUX-mimic is that thesis at frontier scale: We've applied our VAM architecture to the strongest video backbone available today, FLUX 3 from Black Forest Labs, and trained it on data from our own robots and wearables. General-purpose dexterity, running on a single GPU on premises. Because the model already understands world dynamics, it needs far fewer demonstrations to learn a new task. This is game-changing for our mission to deploy robots to factory floors, where industrial robot data is scarce and expensive to collect. We're now testing and deploying FLUX-mimic with manufacturing leaders like @Audi, on complex, multi-step manipulation long considered impossible for conventional automation.
26
97
595
118,891
Introducing FLUX-mimic, a next-generation Video-Action Model for general purpose dexterity, developed in partnership with @bfl_ai. Late last year we published mimic-video and introduced Video-Action Models (VAM): a new family of robotics foundation models built on top of video generation models. We showed that robot control reduces to visual prediction, and that robot capability is downstream of improvements in video modeling accuracy. The obvious implication was that advances in the video modeling frontier would directly translate to increased capabilities in end-to-end robot learning. FLUX-mimic is that thesis at frontier scale: We've applied our VAM architecture to the strongest video backbone available today, FLUX 3 from Black Forest Labs, and trained it on data from our own robots and wearables. General-purpose dexterity, running on a single GPU on premises. Because the model already understands world dynamics, it needs far fewer demonstrations to learn a new task. This is game-changing for our mission to deploy robots to factory floors, where industrial robot data is scarce and expensive to collect. We're now testing and deploying FLUX-mimic with manufacturing leaders like @Audi, on complex, multi-step manipulation long considered impossible for conventional automation.
26
97
595
118,891
From day one, mimic has been focused on a single goal: general-purpose dexterous manipulation. Today we're proud to announce the mimic hand M1 and the mimic wearable U1. We believe the only way to solve dexterous manipulation at scale is by going full-stack at the frontier of physical AI, building every layer ourselves around one fixed point, the human hand. The M1 is a highly backdrivable, tendon-driven hand that covers the full range of human capability, from heavy payloads to fine manipulation.
18
65
513
93,604
The U1 is an exoskeleton that records human demonstrations directly matched to the M1's kinematics, no robot required. Dexterity is one of the last barriers between AI and the physical world. Solving it unlocks entire categories of physical work that automation has never been able to reach.
2
7
79
16,938
This is the foundation on top of which our frontier AI work is done, and we’re just getting started. Full thesis linked here. mimicrobotics.com/blog/solvi…
2
28
1,972
mimic retweeted
A few months ago, I travelled to Zurich together with @NVIDIArobotics to spend time inside one of Europe's most exciting robotics startups... @mimicrobotics! Their mission? To solve one of the hardest problems in robotics: human-level dexterity. Because before robots can replace physical work, they first need to master the human hand. In this episode, we go behind the scenes with the team building AI-powered robotic manipulation, discuss why dexterity is still one of the biggest bottlenecks in automation, and explore what it will take to bring truly capable robots into factories. Mimic Robotics is part of the NVIDIA Inception, a program for startups and VCs, and is leveraging Cosmos to advance AI-powered robot learning. If you're interested in Physical AI, robotic manipulation, and where this industry is actually heading, I think you'll enjoy this one. Timestamps: 0:00 Inside Mimic Robotics 0:24 Building the future of robot manipulation 1:19 The mission behind Mimic Robotics 1:35 The AI powering autonomous robots 2:14 Why Mimic chose NVIDIA Cosmos 2:34 Video foundation models vs. traditional robotics 2:46 Building a multidisciplinary robotics team 3:24 Bridging customers and cutting-edge AI 4:40 Why Zurich is the perfect robotics hub 5:20 Europe vs. America: Scaling a robotics startup 6:30 Bringing generative AI to factory robots 6:49 The race to build Europe's robotics hyperscaler ♻️ Know someone working in robotics? They'll probably want to see this.
6
8
45
6,382
Check out the latest episode of @RoboPapers featuring @elvisnavah, mimic-video and Video Action Models!
Robotics fundamentally involves understanding the dynamics of how things change in the world in response to action and force. This is impossible to learn from static images; instead, it’s far more effective and more data-efficient to learn from video. @elvisnavah joins us to talk about @mimicrobotic. One of the key findings from mimic-video is that pretraining on webscale video allows robots to learn physics priors; as a result, policies train faster, generalize better, and are capable of more impressive dexterity, versus training on static images or image-language pairs as per a VLM. Watch Episode #81 of RoboPapers with @micoolcho and @chris_j_paxton to learn more!
1
10
2,747
mimic retweeted
I love @RoboPapers so it was cool to be invited and chat with @chris_j_paxton and @micoolcho and talk about VAMs and mimic-video!
Robotics fundamentally involves understanding the dynamics of how things change in the world in response to action and force. This is impossible to learn from static images; instead, it’s far more effective and more data-efficient to learn from video. @elvisnavah joins us to talk about @mimicrobotic. One of the key findings from mimic-video is that pretraining on webscale video allows robots to learn physics priors; as a result, policies train faster, generalize better, and are capable of more impressive dexterity, versus training on static images or image-language pairs as per a VLM. Watch Episode #81 of RoboPapers with @micoolcho and @chris_j_paxton to learn more!
2
4
29
4,004
With mimic-video, we were among the very first to propose Video-Action Models for robotics. Today, we are open-sourcing the recipe.
3
34
284
42,605
The @joinplutohouse X @mimicrobotics hacker house in Zurich was 🔥
1
3
12
1,551
mimic retweeted
Seeing the truly insane @mimicrobotics team bring this to life in a very compressed timeline was truly something else. Super proud to be working with @AudiOfficial on the bleeding edge of end to end manipulation!
First look at our collaboration with @AudiOfficial on bringing AI-driven robotics into industrial production. Our end-to-end pixel-to-action model, running on our bi-manual platform, is capable of performing a complex, dexterous and long-horizon insertion task.
6
1
40
5,304
First look at our collaboration with @AudiOfficial on bringing AI-driven robotics into industrial production. Our end-to-end pixel-to-action model, running on our bi-manual platform, is capable of performing a complex, dexterous and long-horizon insertion task.
16
45
314
66,168
mimic retweeted
I just wrote a blog post on mimic-video, @mimicrobotics' answer to a question many have been asking: Why are state of the art VLAs for robotics built on top of Vision-Language Model (VLM) backbones, if those backbones are not pre-trained with physical knowledge in mind?
11
15
160
16,704
Robots might learn better from video than from language! 📼 Most Vision-Language-Action (VLA) models learn what to do from text, but still struggle with how things move in the real world. That makes them data-hungry and slow to train. @mimicrobotics video takes a different route. Instead of grounding robot control in text, it grounds it in video, using large pre-trained video models that already capture physical motion and dynamics. The idea is straightforward: let the video model handle “what will happen next,” and let a smaller control model focus only on turning that visual plan into robot actions. The result is big gains in practice. Robots trained this way need 10× less data, converge twice as fast, and perform better on both simulated benchmarks and real bimanual manipulation tasks. If robots can “imagine” motion using video, control becomes a much simpler problem. Shoutout to Jonas Pai, Liam Achenbach, Oier Mees, @elvisnavah and the rest of the team! Here's the project page: mimic-video.github.io/ ~~ ♻️ Join the weekly robotics newsletter, and never miss any news → ziegler.substack.com
14
45
362
50,122
Today @mimicrobotics and friends are excited to share mimic-video, a new class of Video-Action Model that elevates video model backbones as first class citizens for robot learning!
17
40
328
88,601