Founder of Dream Machines @dm_robots, building ML models for robotic arms. Previous: Stats at @ETH & @NeurIPSconf. I post about RL and VLA post-training.

Zurich, Switzerland
On April 2nd, I incorporated Dream Machines and raised an angel round. This is the document I raised it on. Most company stories are told post hoc and cleaned up. I want to share ours ex ante. Europe needs more technical people building what they see the world lacking, and that’s easier when you can see what a start actually looks like for others. Since writing this memo, the first founding engineer has joined, more companies have sent us their products, and we have a first model running that we’ll deploy at our first customer shortly. We're looking for exceptional ML engineers/researchers to join our mission as founding engineers and research interns. If our view resonates with you - or you have a name of somebody we should be talking to - reach out. DM's are open. Link to memo below.
22
3
193
26,577
Great piece. Now let's run the numbers for Europe. > Germany: same count ... but with a quarter of the population. > Europe: 8x > Robot-to-worker density: 2x North America and rising. The US market is cool, but the big opportunity nobody's talking about is the Mittelstand.
America has 16,876 machine shops. 83% employ fewer than 20 people. Just nine employ more than 500. New defense companies are moving from prototypes to mass production, and every missile, drone, ship, and rocket they ultimately produce needs to run through one of these small shops. That friction is a national security liability. Connor Love and Collen Larson on the next generation of American manufacturing that gets built by smoothing it: a16z.news/p/the-case-for-the…
11
4
70
7,676
Dominique Paul retweeted
Did you ever see someone flipping the sign of your loss, and achieve great results? Well, this summer, we did. Some papers (refs below) suggested to reverse the self-distillation loss, i.e, move away from the teacher. We wrote a blogpost on our investigation and our findings.
Self-distillation with privileged context shifts behavior and can hurt reasoning. Repulsion reverses the shift but becomes unstable. In our experiments, contrasting both signals improves reasoning while keeping training stable. Blog: self-distillation.github.io/…
1
4
14
966
Kudos to @raminthies for extending the Focus Feed Chrome extension to Instagram!
I open X and LinkedIn to reply to messages or post an idea, and too often the feed pulls me in when I'm in focus mode. So I built a Chrome extension that fixes that. > Hides the feed and notifications. Messages and posting still work > Unlocking takes 15 seconds, and switching tabs resets the timer. Long enough for the better part of your brain to change its mind > Hiding it again is one click > Every morning at 6 AM it's all hidden again, so you don't have to remember github.com/DominiquePaul/foc…
3
746
I love evals - said no roboticist ever.
7
3
152
6,860
Partly. Monocultures are also productivity machines. When everyone around you already accepts the premises, you stop burning energy defending them and can push an idea much further than you could anywhere else. Plenty of extreme bets only get made when nobody around you makes you question the premises. Doesn't make the post wrong: You want one foot inside the monoculture for the leverage, and one foot outside so it doesn't decide what you want.
This is why Peter Thiel was so successful. One of his biggest inspirations was the French philosopher René Girard for his work on mimetic desire (and it's consequences) - here's how it links to @shaunmmaguire's points: "I left Silicon Valley... Because it's a monoculture... The most extreme group think of any place on the planet." Girard's central thesis was basically: humans don't desire things autonomously, they copy the desire of others. People fiercely defend the belief that their desires are individual/authentic (according to Girard "romantic illusion") when in reality they subtly train each other to want the exact same thing (in this particular monoculture it would be status associated with being a tech founder - ARR metrics, raise sizes, valuations, "frontier" ideas etc). "So what do they do? They start stack-ranking each other... founder of co $100B market cap > founder of co with $50B market cap..." Girard believed that when actors are in close status, they become mimetic rivals. Everyone is competing for the same narrow definition of success within the same tight ecosystem. They end up trapped in an ordered hierarchy where minute differences in status ($100B vs. $50B drive some sort of psychological rivalry). "Pure toxicity that nerfs creativity... Reject this nonsense. Stay weird. Think independently." Girard believes that intense mimetic competition creates homogeneity, sheep like or heard like behaviour. Instead of thinking about how can I create novel technology (a Facebook, a SpaceX, an AirBnB) new entrants focus on "how do i beet my peers at the game?" This was probably the most dangerous part of a monoculture made up of mimetic rivals (from an investors point of view) - because theres no real alpha left - the competition just competes it all away. Everyones fighting over the same piece of the pie. Where do you think he got his "competition is for losers" line from? He got it from Girard. Founders Fund was basically Thiel's vehicle to invest in anti-memetic ideas that Silicon Valley actively ignored. 1. Palantir - defence was taboo + governments were the anti-thesis of any start-ups philosophy 2. SapceX: Manufacturing and CapEx intense businesses were to be avoided - light scalable software was better 3. FB: Literally a real world social graph built around mimetic cultures (here i guess he was like ok we can profit of mimetic cultures lol). Everyone should go study René Girard and then go look at all of Thiel's conference talks - it's the same lines over and over again. pretty cool! p.s. this is not exclusive to SF - it's exclusive to humans
2
1
38
6,443
Dominique Paul retweeted
First experiments on metadata conditioning failed completely :/ @physical_int adds different metadata about the episode or action chunk to the text prompt for their π0.7 model. I wanted to try this with speed. I took a dataset consisting of two recording sessions where one session had episodes with double the speed of the other. I added a suffix to the promot: “speed: fast|slow” Unfortunately the policy completely ignored the text prompt and the behaviour didn’t change when I switched “slow” and “fast”. After analysing the datasets I found that the information about the speed could already be extracted from the images themselves. A logistic regression was able to tell the two sessions apart on 99.8% of held out frames. Looking at the data the only difference I can see is slightly different placement of the box and tray on the table. but I guess this is enough. Learning: If a conditioning signal is recoverable from the observation, it will simply be read out from it and the text is ignored.
1
12
3,392
Dominique Paul retweeted
HG-DAgger can lead to a worse policy if done wrong. After the second iteration of intervention data collection I observed that the policy (π2) barely improved on success rate and even had a much longer mean episode duration. Reason for this is the quality of the interventions. It is hard to correct a policy from a bad situation. Your movements are slower and less smooth than normal demonstrations because you are not in the flow. Too many of such interventions teaches the policy to act slower and generally needs more recoveries. I found a quick fix for this though: finish the training run during the annealing with clean demonstrations only. This mostly recovers the faster and more precise movement from the demonstrations while keeping the skills to recover from the interventions. This brought the success rate on our task (full box) up to a 88% success rate.
4
6
36
2,617
There are so many robotics companies, each with its own approach, and trying out a new tool is a lot of friction for all of them. That makes trust worth a lot: Companies building tools for robotics might take a while to land their first customers, but once a tool visibly works for a few teams, others will pile in fast.
1
32
1,596
Day 2 of working on RLT: >Noise injection off. The paper adds fixed Gaussian noise but doesn't say how much. The jitter stayed anyway, and most of it came from chunk handovers: the RLT actor edits each chunk by a few degrees, then the next VLA chunk snaps back to the VLA's own path. Added a 5-step (100 ms) fade at each handover, which should hopefully fix it (to be tested tomorrow). > Regularisation towards the original policy up 5x (0.05 → 0.25) to keep edits closer to the VLA. Probably better to start low and ramp it up once the actor has learnt something (not doing that yet). The right value most likely differs per task. > Added a 1,500-transition warm-up (~5 min of RLT-active time). Yesterday I tried without it, but that caused bad behaviour after 5 rollouts. Better now, but with RLT on it still performs worse than the base policy: 72% success during warm-up vs 15% with the actor, counting takeovers as failures. > RLT is only active for part of each trial, so it took 25–30 min of rollouts (46 trials) before the actor drove at all. This slows down feedback on which hyperparameters work well. > 3x512 MLP. The paper uses 2x256 for three tasks and 3x512 for the hardest (screw installation). Their critical phase starts from lightly randomised setups. Our object positions vary much more. > Now logging RLT metrics to W&B (only logged RL token training before, not live updates). Judging hyperparameters from rollouts alone is really hard. > One thing I was thinking about: replay is uniform, so the critic spends most of its time on older data. Recency weighting could help. On the bright side: we have a customer demo next week and need to collect some more DAgger data for a good demo (not a sustainable solution, but we don't want any major fuckups at the demo). RLT leads to more OOD states, so it's at least a convenient setup for collecting intervention data. It's great that @physical_int released the paper, but most of the hyperparameters you need aren't in it (noise, regularisation, batch size, LR, discount, buffer size), so it's super hard to get RLT working. Will continue to post updates.
4
1
83
3,516
Dominique Paul retweeted
Good writeup that lines up with my own experience for this kind of thing. Much less data, more diverse, higher quality.
Two months of ablating π0.5 finetunes on a real manufacturing task. The policy now succeeds 98% of the time. We're publishing all results, all the data, and every run, including the ones that went nowhere. What we learned: 1/ 5x more data took us from 63% to 76%. Pure scaling was the weakest lever we found. 2/ Diversity got more out of the same hours. 4h spread across five scenes beat 4h in the eval scene by 30pp. 3/ Quality did even better. One clean, slower hour on top of the 21h: 76% → 90%. That single hour added more than the 17 hours before it. 4/ So we tried quality on its own: 1h of clean data plus 240 rollouts with human interventions. 1.7h in total, 28% → 88%. That's 12 points above the 21h model with a twelfth of the data. 5/ Inference settings add more on top, with no retraining. They took the 21h model from 76% to 93%, and the 21h+1h model from 90% to 98%. Just as useful: what didn't matter. Learning rate, batch size, relative vs. absolute joints, image augmentation. All less important than we expected. Full write-up linked below.
1
4
63
7,475
Working on first real-world RL token experiments today. > I’m only activating it for the critical part (insertion) as in paper > first rollouts with RLT work decent > After ~10 rollouts with updates the policy gets very jerky - maybe because I’m injecting noise for exploration or the residual policy isn’t regularised strongly enough. Noise injection should show up as jerkiness is first rollouts as well though? In the video you can see the jerkiness. Also, once I turn the RLT network off and unedited VLA takes over everything runs smooth and actuator is inserted cleanly.
4
6
106
5,507
You can see the progression of the success rate here across rollouts.
1
5
732
Jerkiness issue before was due to noise injection. Previous noise was sigma=2.5% of each joints range which was way too much. Set it to a tenth of that and it works better. Still very slow to learn and does stupid things, but you can tell there‘s a learning effect
1
3
516
Dominique Paul retweeted
Stoked to be part of this team! 🙌
Brainstorming has begun… In the next 9 months we, a team of 9 mechanical engineers from ETH Zurich, are building a drone that autonomously intercepts consumer drones nondestructively. It secures the target and returns it safely. Looking forward to sharing the journey!
1
7
719
Dominique Paul retweeted
ETH Zurich’s Junzhe He demonstrates a humanoid robot playing badminton against a human.
6
13
115
8,636
.@JanneStiefel is taking a year off from ETH to build an X-wing style drone that snatches other drones out of the sky. He's also joining us at Dream Machines part-time to work on our next grippers and hardware - open source oc.
Experimented with different TPU95 inserts shapes for our grippers at @dm_robots this week. Favorite so far: the vertical/tilted ribs, the TPU bends to the middle when compressed, curls almost like a finger. The Honeycomb turned out too stiff, still some work to do there…
5
2
98
7,571
I open X and LinkedIn to reply to messages or post an idea, and too often the feed pulls me in when I'm in focus mode. So I built a Chrome extension that fixes that. > Hides the feed and notifications. Messages and posting still work > Unlocking takes 15 seconds, and switching tabs resets the timer. Long enough for the better part of your brain to change its mind > Hiding it again is one click > Every morning at 6 AM it's all hidden again, so you don't have to remember github.com/DominiquePaul/foc…
8
1
44
3,148