We should stop optimizing robot policies against a single overall reward. Trajectories differ along many axes, such as speed, precision, and subtask completion, and one can be better on some while worse on others. If we collapse all of that into a single overall axis we lose this structure making the reward ambiguous and harder to optimize. Blog: freeform-pl.github.io/fpl.we… Paper: arxiv.org/abs/2606.32027
10
32
161
34,932
Marcel Torné retweeted
AI can increasingly make progress on open problems like Navier-Stokes. Yet defining a new problem still relies on human research taste. Introducing ScholarCatalyst, a benchmark built from AI researchers’ firsthand accounts of what inspired their work.
5
44
209
20,112
Marcel Torné retweeted
Great researchers have an uncanny ability to make connections that seem inevitable in hindsight, in places nobody else would have thought to look. Can we measure this ability? Introducing ScholarCatalyst: a far-from-saturated benchmark for finding what we call "catalyst papers"📚, labeled by 184 lead authors on 207 of their own recent projects. Paper: arxiv.org/abs/2610.02202 To make sustained progress on open-ended problems, I think agents need the sort of "research taste" that great researchers have. They need to make deep connections between earlier discoveries and problems those discoveries weren't intended to solve. ScholarCatalyst is a first step towards this goal. More details in the thread below🧵
16
49
305
21,214
Marcel Torné retweeted
There’s lots of buzz around agentic harnesses for robots, particularly for long horizon tasks that require complex reasoning and memory. But what will it really take to turn a reasoning agent into a reactive, reliable and low-latency robot policy? In our new paper, workspace models, we design a new memory architecture that acts as a latent harness for stronger reasoning models. Here’s why we think this might be the way forward: (1/9)
7
54
200
28,866
Marcel Torné retweeted
Introducing Joga, a full-stack humanoid soccer system with active vision. Agile dribbling. Tight direction changes. Receive → dribble → shoot/pass. An actuated neck enables ball tracking with onboard vision. Residual models for perception + actuation close the sim-to-real gap. Paper coming soon! Work done in collaboration with @Avi_Narula14 @khai_ngx @venki1899 @JohnMarangola @martinpeticco @pulkitology
23
39
268
55,841
Marcel Torné retweeted
Specifying reward functions for robots is one of the hardest things about reinforcement learning. Robot rewards often need to be very detailed; metrics like progress can be ill-defined and hard to estimate. This leads to most robot learning defaulting to sparse rewards or simple preference learning. But naive preference learning (having human annotators choose one trajectory over another) is an easy solution, but obscures a lot of the signal in complex tasks and can make learning a lot less efficient. @marceltornev, @anubhamahajan01, and @AbhijnyaBhat join us to talk about their solution: freeform preference learning, which lets annotators define natural-language axes to compare trajectories over. This improves real-world performance on long-horizon manipulation tasks over sparse rewards and simple binary preference learning. Watch Epsiode 103 of RoboPapers, with @micoolcho and @DJiafei, today to learn more!
2
3
21
17,476
We presented Freeform Preference Learning at the @RoboPapers podcast! Check it out! Paper: arxiv.org/abs/2606.32027
Specifying reward functions for robots is one of the hardest things about reinforcement learning. Robot rewards often need to be very detailed; metrics like progress can be ill-defined and hard to estimate. This leads to most robot learning defaulting to sparse rewards or simple preference learning. But naive preference learning (having human annotators choose one trajectory over another) is an easy solution, but obscures a lot of the signal in complex tasks and can make learning a lot less efficient. @marceltornev, @anubhamahajan01, and @AbhijnyaBhat join us to talk about their solution: freeform preference learning, which lets annotators define natural-language axes to compare trajectories over. This improves real-world performance on long-horizon manipulation tasks over sparse rewards and simple binary preference learning. Watch Epsiode 103 of RoboPapers, with @micoolcho and @DJiafei, today to learn more!
1
3
18
5,782
Marcel Torné retweeted
Full episode dropping soon! Geeking out with @marceltornev @anubhamahajan01 @AbhijnyaBhat on Freeform Preference Learning for Robotic Manipulation freeform-pl.github.io/fpl.we… Co-hosted by @micoolcho @DJiafei
4
8
1,477
Had lots of fun presenting MEM (pi.website/memory) at @ycombinator Paper Club and meeting lots of people excited and working on robotics. Thank you @FrancoisChauba1 for the invitation!
This week’s Paper Club is all about robotics. Every year for the last decade, someone has promised that the era of robotics is just around the corner. But we’re still waiting. So we gathered a bunch of the top researchers working in AI and robotics to present the latest findings on where we are and what comes next. Thanks to the following presenters: 0:00 – @FrancoisChauba1: Ten years of “next year, robotics is solved” 7:59 – @marceltornev: MEM - Multi-Scale Embodied Memory for Vision Language Action Models (arxiv.org/abs/2603.03596) 20:21 – Milan Ganai: Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning (arxiv.org/abs/2602.08167) 33:42 – @tylerlum23: SimToolReal - An Object-Centric Policy for Zero-Shot Dexterous Tool Manipulation (arxiv.org/abs/2602.16863) 51:21 – @NikolausWest: Why the next great robotics companies will start with teleoperation 1:08:30 – @BillJiao930 & @Guanming717: World action models and what comes after VLAs
3
2
33
13,139
Marcel Torné retweeted
Catching skin cancer early is a home robotics problem. Melanoma is highly treatable when detected early, yet today’s screening process depends heavily on patients noticing tiny changes across their entire skin surface. This requires patients to solve a near-impossible visual-memory and registration problem. I built OpenDerm, an open-source 4-DOF robot that captures high-resolution images of the skin and uses them to reconstruct and track the skin surface in 3D over time. The best way to make skin screening truly routine is to bring it into the home. OpenDerm shows that inexpensive robotic skin imaging is possible, but the path to scale is not a dedicated screening robot in every household—it is to make skin screening one of the many useful things a general-purpose home robot can do. Read more about why I built OpenDerm and how it works here: Blog: marionlepert.github.io/blog/… Project: openderm.github.io/
416
977
9,063
1,505,772
It's very impressive to see how the Real-to-Sim-to-Real pipelines have improved since our results in 2024 with RialTo! real-to-sim-to-real.github.i… Congrats to the team @SceniXai @theworldlabs for all of the advancements! Looking forward to what's coming next.
Replying to @drfeifei
The Real-to-Sim part of R2S2R transforms physical robots, sensors, environments, objects, and interactions into simulations that preserve task-relevant observations and dynamics - not only how the world looks, but how it acts when the robot interacts with it. These results set a new bar for sim-real alignment in robot manipulation.
1
2
38
5,000
I will be giving a talk at #RSS2026 workshop on "Towards Robust Execution of Long-Horizon Whole-Body Central Tasks" at 11:30 am Sydney time. I'll be talking about our work done at @physical_int MEM: Multi-Scale Embodied Memory for Visual Language Action Models. Check the thread below for more details on our work! I will unfortunately be remote this time but hope to see you all soon at the next conference!
We equipped PI policies with memory! And taught our robots to do long-horizon real world tasks such as preparing the items for a recipe, cooking a grilled cheese and cleaning the kitchen!
1
3
35
13,041
Marcel Torné retweeted
I unfortunately had to cancel my #RSS2026 trip last minute, but fortunately my excellent students, collaborators and postdocs will be representing our work much better than me anyways :) We have 4 papers that you might enjoy (Sydney time): 1. Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation sim-dist.github.io/ (Mon 7/13 3:15-4:00pm) @ty_westenbroek @jakejlevy 2. TMRL: Diffusion Timestep-Modulated Pre-training Enables Exploration for Efficient Policy Fine-tuning weirdlabuw.github.io/tmrl/ (Thu 7/16 3:15-4:00pm) @matthewh6_ @Jesse_Y_Zhang 3. PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies polaris-evals.github.io/ (Tue 7/14, 11:50-12:30pm) @prodarhan @KarlPertsch 4. Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparison robometer.github.io/ (Wed 7/15, 3:15-4:00pm) @Jesse_Y_Zhang @aliangdw @yigitkkorkmaz Please go ask them many difficult questions! :) Also - go see the amazing @Jesse_Y_Zhang at his RSS pioneers poster Tue 4-5pm!
1
20
157
11,065
Marcel Torné retweeted
I'm giving a talk on how we can move beyond the scalar reward bottleneck for both robotics & LLMs. ICML RLxF workshop tomorrow at 1:30 pm.
RL is hitting a ceiling with human feedback. What if the world itself becomes the signal? Join us at the RLxF: RL from World Feedback 🌍 workshop at ICML 2026 @icmlconf tomorrow (July 10th)! Web page: sites.google.com/view/rlxf-i…
9
21
274
54,113
Marcel Torné retweeted
We’ve been approaching reward supervision for robots the wrong way. I think freeform preferences are part of the answer. A short 🧵
6
41
395
37,154
Marcel Torné retweeted
What if the way we collect human feedback in robotics is quietly losing information? If one trajectory drops the cutlery and another drops the plate, asking for a single preference hides information behind the choice. Freeform Preference Learning lets annotators describe preference axes in natural language, learns rewards conditioned on those axes, and extracts better policies. Congrats to @marceltornev, @anubhamahajan01,@AbhijnyaBhat, and @chelseabfinn!
We should stop optimizing robot policies against a single overall reward. Trajectories differ along many axes, such as speed, precision, and subtask completion, and one can be better on some while worse on others. If we collapse all of that into a single overall axis we lose this structure making the reward ambiguous and harder to optimize. Blog: freeform-pl.github.io/fpl.we… Paper: arxiv.org/abs/2606.32027
5
6
59
13,583
Marcel Torné retweeted
Our intern just built the first zero-person company. Listen's agent ran a loop: - Interview users - Build - Test with real people - Fix issues - Repeat 2,000 interviews and 100 concepts later: an app with 100s of paying customers. Here’s how it works:
101
68
967
587,787
We should stop optimizing robot policies against a single overall reward. Trajectories differ along many axes, such as speed, precision, and subtask completion, and one can be better on some while worse on others. If we collapse all of that into a single overall axis we lose this structure making the reward ambiguous and harder to optimize. Blog: freeform-pl.github.io/fpl.we… Paper: arxiv.org/abs/2606.32027
10
32
161
34,932
We also ran a qualitative analysis of the learned reward model. Reward models learned from freeform preferences show better credit assignment than those learned from binary preferences. They seem to identify relevant events more accurately without any explicit task segmentation.
1
1
8
467
This work wouldn't have been possible without my collaborators @anubhamahajan01 and @AbhijnyaBhat. And the great advising from @chelseabfinn. Thank you all! 🙏 Our blog post and code are available at: freeform-pl.github.io/fpl.we… And the paper at: arxiv.org/abs/2606.32027
1
1
11
439