Asst Prof of CS & EE @Stanford Co-founder of Physical Intelligence @physical_int PhD from @Berkeley_EECS, EECS BS from @MIT

Palo Alto, CA
LLM post-training used to mean fine-tuning to a downstream task Robotics has been stuck in this setting, needing task-specific fine-tuning for best performance π07 changes this: It works out of the box & outperforms fine-tuned specialists Details: pi.website/pi07
31
66
631
79,090
New blog post with @perryadong on what we need for robots to be broadly useful in the real world. pd-perry.github.io/posts/pos… Reliability is the one of the biggest open challenges in AI right now. Current models work out okay if a person will be reviewing the outputs (eg drafting code), but it will become more of a bottleneck as we want systems to act with more autonomy and more trust.
What will be the “RLHF” moment for robotics? What will it take to get robotics to where LLMs are today and beyond? New blog post with @chelseabfinn sharing some thoughts on the state of RL for frontier robotics models and what's missing 👇 Blog: pd-perry.github.io/posts/pos…
17
53
450
45,701
RL for dynamic tasks requires handling latency. Building on EXPO-FT, we let the base VLA operate on older images while a small edit policy operates in real time. Outperforms RTC by a significant margin, across the board. Paper: pd-perry.github.io/real-time…
Introducing Real-Time EXPO-FT – Fast and Reliable RL for Real-Time VLA Policies! Real-Time EXPO-FT unlocks π0.5 on challenging dynamic tasks, such as balancing a ball on a plate and striking a ball into the goal (1/6)
9
23
225
26,520
Project led by @perryadong and @khhung906 with @DorsaSadigh at @StanfordAILab. Plus some related recent work from friends at Berkeley/Siemens:
Asynchronous VLA inference reduces inference delay, but breaks the Markovian assumption necessary for RL fine-tuning. How can we enable RL fine-tuning of VLAs with async inference? We introduce ARLI: Asynchronous RL with Intermediate Information! async-rl-intermediate-inform… (1/n)
2
4
2,915
A video from a Pi robot deployed at Dandelion Chocolate, fully autonomous w/ no interventions. 🤖 Deploying robots has taught us surprising lessons about the gap between proof-of-concept (i.e. building one box) and real-world utility (productively building boxes for hours). I expected that the hard part is building the box, since it’s the most dexterous, but that wasn't what we found. Counterintuitively, the hardest part was reliably stacking the boxes. Our original table-mounted robot had poor visibility of the stack without special separately-mounted cameras. Plus, stacking requires more generalization (each box is placed in a different location), and an imprecisely-placed box can lead the entire stack to collapse many boxes later. We recently switched this deployment over to a mobile robot, and it's fun to watch the robot being completely self sufficient for multiple hours. 🙂 Data and feedback from real world use-cases like these are quite valuable for the π pre-trained model as we scale!
46
67
918
104,133
Congrats to @zipengfu, @chenwang_j, and co on the release!
Introducing OM-1, our first robot foundation model, zero-shot generalizing to any robot: table-top arms, industrial arms and humanoids. - learned directly from human manipulation data - no teleop/robot data - close to human-level dexterity and efficiency - multi-robot collab
17
32
542
76,095
Chelsea Finn retweeted
Harness optimization is sample-efficient but plateaus. What should you do if you can afford to update the model too? Introducing WHALE: a simple recipe for jointly optimizing an LLM's weights and harness. Blog: krafton.ai/blog/whale/ Paper: arxiv.org/abs/2609.00196
5
64
469
44,853
One of the most important aspects of scientific discovery is deciding where to draw insights from. While LLMs are promising tools for science, we lack datasets & evaluations for this step. Help contribute to a public dataset for exactly this: tinyurl.com/45b9ykae
We’re asking the research community to help us build a benchmark for research taste. Scientific discovery starts with a fundamental step: which prior work is worth building on? We want to capture this undocumented layer through our collective knowledge. Please sign up: tinyurl.com/45b9ykae ↓
9
24
271
54,042
Chelsea Finn retweeted
World models have emerged as one of the biggest directions in physical AI. At the same time, RL fine-tuning is unlocking capabilities in frontier models beyond what pretraining can achieve on its own Can we get the best of both worlds? We propose Q-Learning with World Models (QWM) (1/7)
7
56
515
32,195
Chelsea Finn retweeted
Robots can already fold laundry, make espresso, clean kitchens, and assemble things. The harder problem is getting them to do those tasks reliably, for long periods of time, without a human babysitting them. At Startup School 2026, @physical_int cofounder @chelseabfinn explains what it takes to build general-purpose robots that work in the real world. She shares how reinforcement learning pushed robot throughput up 2x, how their systems can run autonomously for hours, and why she believes robotics is entering its GPT era: moving from specialized models toward general-purpose systems that can work across tasks, robots, and environments. 00:00 — The State of Physical Intelligence 01:23 — What It Takes to Make Robots Useful 05:11 — The Reliability Problem 07:43 — Reinforcement Learning for Robotics 09:35 — Learning From Failures 12:43 — Training Robots to Improve Themselves 14:21 — Can a Robot Work for 13 Hours Straight? 17:36 — Why Robots Need Memory 21:22 — Building a General-Purpose Robot 25:02 — From Fine-Tuning to Out-of-the-Box Models 27:35 — Training on All the Data 30:20 — One Model That Beats the Specialists 31:21 — Compositional Generalization 37:49 — The GPT Era of Robotics 39:49 — Q&A
37
45
320
113,098
Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch. We find that pretraining Q-functions on data from diverse policies is critical to see improvements from pretraining. Paper: arxiv.org/abs/2607.27203
Pretraining has worked remarkably well across domains We show this doesn’t hold for Q-functions in online RL from a pretrained policy — and propose IPE, a more effective way to learn Q-functions for online RL fine-tuning (1/6)
13
54
642
79,568
I'm giving a talk tomorrow at ICML on emergent physical generalization, including π0.7 🤖 3:15 pm @ SCALE workshop in Ballroom 201 scale-icml-2026.github.io/
12
20
368
28,558
I'm giving a talk on how we can move beyond the scalar reward bottleneck for both robotics & LLMs. ICML RLxF workshop tomorrow at 1:30 pm.
RL is hitting a ceiling with human feedback. What if the world itself becomes the signal? Join us at the RLxF: RL from World Feedback 🌍 workshop at ICML 2026 @icmlconf tomorrow (July 10th)! Web page: sites.google.com/view/rlxf-i…
9
21
275
54,108
Freeform preference learning has multiple nice properties: (1) It works better When controlling for the number of preference queries, learning with multi-axis preferences yields far more performant policies than single-axis rewards.
2
1
19
4,336
(3) Long-horizon credit assignment Most robot RL focuses on short horizon tasks b/c dense temporal rewards are hard to get. Freeform preferences yield dense rewards for subtasks without subtask segmentation.
2
1
16
3,939
Project led by @marceltornev, @anubhamahajan01, @AbhijnyaBhat Paper: arxiv.org/abs/2606.32027 Code & videos: freeform-pl.github.io/fpl.we… Check out Marcel’s thread for more details!
We should stop optimizing robot policies against a single overall reward. Trajectories differ along many axes, such as speed, precision, and subtask completion, and one can be better on some while worse on others. If we collapse all of that into a single overall axis we lose this structure making the reward ambiguous and harder to optimize. Blog: freeform-pl.github.io/fpl.we… Paper: arxiv.org/abs/2606.32027
5
1
35
5,830
Freeform preferences let the supervisor define relevant axes and then specify preferences along those axes. Axes can be either a fixed rubric or freeform language. This eliminates ambiguity, allows for thorough coverage of all axes, and provides more dense supervision.
1
1
15
3,346
With freeform preference data, we train a lang-conditioned reward that captures all axes of a task. We then train a policy conditioned on each reward axis and the corresponding reward. FPL allows the robot to maximally leverage and learn from each axis of supervision.
1
1
10
3,102
Sparse rewards, progress metrics, and preferences are popular, but they - often neglect many aspects of a task - collapse many axes into one measure - frequently yield ambiguity and disagreement across annotators We instead propose freeform preference learning
1
14
4,025
We’ve been approaching reward supervision for robots the wrong way. I think freeform preferences are part of the answer. A short 🧵
6
41
395
37,152
Long term, we need reward models to capture all aspects of performance, like success, outcome quality, and speed. This also includes task-specific axes like: - was the PB spread evenly? - was the apple slightly bruised while bagging it? - was the furniture bumped or scratched?
1
24
5,295