Learning and planning for safe, embodied autonomous systems under uncertainty. Senior Research Scientist @ToyotaResearch. PhD from @StanfordMSL. 日本語 & English

California, USA
Haruki Nishimura retweeted
A new negative-result investigation: closer-to-target (non-Gaussian) action priors did not improve the fine-tuning of Large Behavior Models (LBMs), i.e., robot policies pretrained on large, diverse datasets. Prior work found that replacing the Gaussian prior of diffusion and flow-matching policies with a closer-to-target prior substantially improves policies trained from scratch. We hypothesized that this benefit would carry over to fine-tuning pretrained LBMs. Across 100K+ simulation rollouts and 1250 hardware rollouts, we found that it does not. Website: cxu-tri.github.io/non_gaussi… Paper: arxiv.org/abs/2609.27070 🧵 (1/10)
2
12
37
4,035
Haruki Nishimura retweeted
(1/12) How should we rigorously compare robot policies? Comparison is central to robotics research, but is inherently expensive. We introduce NSCORE, a flexible, general-purpose, and data-efficient method for rigorous policy comparison. Accepted to RSS 2026.
1
6
27
3,048
Sample-efficient and reliable policy comparison is essential for both fast design iteration and reproducible baseline comparison, where expensive hardware eval remains the gold standard. Check out this awesome RSS work led by @das_princeton on a new, general framework!
(1/12) How should we rigorously compare robot policies? Comparison is central to robotics research, but is inherently expensive. We introduce NSCORE, a flexible, general-purpose, and data-efficient method for rigorous policy comparison. Accepted to RSS 2026.
7
1,099
Haruki Nishimura retweeted
How should we evaluate robots in the age of foundation models? We hosted the RSS RoboEval Workshop with folks from academia, industry, and policy to discuss this. 💪 We put together an actionable guide and insights to get you started on robot evaluations. Link below. 🧵1/N
1
10
24
2,746
Robot evaluation is an open problem, especially in the age of foundation models. Check out this blog post sharing actionable insights from our RSS Workshop last year, highlighting challenges and describing best practices. Huge thanks to @hocherie1 for leading this effort!
How should we evaluate robots in the age of foundation models? We hosted the RSS RoboEval Workshop with folks from academia, industry, and policy to discuss this. 💪 We put together an actionable guide and insights to get you started on robot evaluations. Link below. 🧵1/N
3
4
463
Haruki Nishimura retweeted
We extended the deadline to *June 22nd*! Submit your coolest and craziest (in-progress or completed) works on generalist robot safety 😎 Workshop co-organized with the great team: @ArpitBahety @kensukenk @imp_aa @RobobertoMM @ianabraha @LihanZha
Excited to announce our #RSS2026 workshop: "Rethinking What It Means to be 'Safe' for Generalist Robots"! 🛡️🤖 Have new work or videos of robot safety failures? Submit by June 12! 👇
1
5
12
2,100
Haruki Nishimura retweeted
What does it actually mean for a modern robot to be safe? As generalist robots move across tasks, environments, and users, safety must encompass many dimensions: collisions, semantic constraints, perceptions of safety, privacy, and more. Join our discussion at RSS 2026!
Excited to announce our #RSS2026 workshop: "Rethinking What It Means to be 'Safe' for Generalist Robots"! 🛡️🤖 Have new work or videos of robot safety failures? Submit by June 12! 👇
4
6
1,761
Haruki Nishimura retweeted
Excited to announce our #RSS2026 workshop: "Rethinking What It Means to be 'Safe' for Generalist Robots"! 🛡️🤖 Have new work or videos of robot safety failures? Submit by June 12! 👇
1
4
14
5,991
Haruki Nishimura retweeted
Releasing RecGen: a collaboration between @ToyotaResearch, @toyota_europe, and @UvA_Amsterdam tackling a core 3D vision challenge: reconstructing complete multi-object scenes (parts, poses, textures, even occluded geometry) from just 1 to a few RGB-D views. Trained purely on synthetic data, RecGen achieves SOTA on real-world robotics and 6D pose benchmarks, handling occlusions, symmetry, and complex interactions. A step toward scalable, high-fidelity digital twins for robotics, and better evaluation and training of generalist policies. reconstruction-by-generation…
3
35
223
27,537
Haruki Nishimura retweeted
I was thrilled to be back at @MIT for the Robotics Seminar! The talk recording is available now: Rethinking Robot Safety & Alignment in the Era of Generalist Policies piped.video/pZM8sgLAye0?si=GG7t…
5
65
9,397
Haruki Nishimura retweeted
A few interesting rollouts from the Foundry-QwenVLA-2.5B multi-task model on seen tasks in sim –  a 🧵. I really like behaviors that involve non-prehensile manipulation, like the little nudges in StoreCerealBoxUnderShelf.
Releasing VLA Foundry: an open-source framework that unifies LLM, VLM, and VLA training in a single codebase. End-to-end control from language pretraining to action-expert fine-tuning — no more stitching together incompatible repos.
2
20
117
14,968
Haruki Nishimura retweeted
Having control over upstream LLM/VLM training is key to training a good robotics model. We hope VLA Foundry opens the door for researchers and practitioners to answer questions they previously wouldn’t even have thought of asking if upstream pretraining was simply inherited!
Releasing VLA Foundry: an open-source framework that unifies LLM, VLM, and VLA training in a single codebase. End-to-end control from language pretraining to action-expert fine-tuning — no more stitching together incompatible repos.
2
6
28
3,585
Haruki Nishimura retweeted
TRIで最後に関わったプロジェクトである、VLA Foundryがついにリリースされました!異なる言語モデルやビジョンモデルを手軽に試せるだけでなく、Drake + Blenderを用いたシミュレーション環境で複数タスクの評価も簡単に行えます。ぜひ試してみてください!
Releasing VLA Foundry: an open-source framework that unifies LLM, VLM, and VLA training in a single codebase. End-to-end control from language pretraining to action-expert fine-tuning — no more stitching together incompatible repos.
17
117
15,308
This is hugely based on @das_princeton's implementation that came out of the collaboration between TLU tri.global/trustworthy-learn… and @Majumdar_Ani's group at Princeton out of an internship project!
This is actually a pretty big deal — we rely on @imp_aa’s implementations to tell when policies are statistically different than each other. If someone presents some quick mean-only results internally without the CLD analysis, you can be sure someone will eventually ask for it.
1
5
835
Haruki Nishimura retweeted
Releasing VLA Foundry: an open-source framework that unifies LLM, VLM, and VLA training in a single codebase. End-to-end control from language pretraining to action-expert fine-tuning — no more stitching together incompatible repos.
10
76
493
75,690
A huge shout-out to TRI's VLA team for the public release of VLA Foundry! You can take full control of VLA training with this fully open-sourced codebase, which comes with a nice GUI dashboard with rigorous policy comparison powered by STEP🪜 tri-ml.github.io/step/
Releasing VLA Foundry: an open-source framework that unifies LLM, VLM, and VLA training in a single codebase. End-to-end control from language pretraining to action-expert fine-tuning — no more stitching together incompatible repos.
11
4
44
7,816
Haruki Nishimura retweeted
Great to see @LeRobotHF using STEP as a tool for statistically rigorous policy comparison! arxiv.org/abs/2503.10966
Releasing the Unfolding Robotics blog! Time to unfold robotics: we trained a robot to fold clothes using 8 bimanual setups, 100+ hours of demonstrations, and 5k+ GPU hours. Flashy robot demos are everywhere. But you rarely see the real story: the data, the failures, the engineering. We’re sharing everything: code, data, and details in the blog → huggingface.co/spaces/lerobo…
3
36
6,401
Congrats to the @LeRobotHF team on this remarkable contribution to the robotics community by open-sourcing "everything" including code, data, and all the valuable knowledge! Our TLU team at TRI is fortunate to have collaborated on statistical evaluation and analysis.
Releasing the Unfolding Robotics blog! Time to unfold robotics: we trained a robot to fold clothes using 8 bimanual setups, 100+ hours of demonstrations, and 5k+ GPU hours. Flashy robot demos are everywhere. But you rarely see the real story: the data, the failures, the engineering. We’re sharing everything: code, data, and details in the blog → huggingface.co/spaces/lerobo…
1
2
9
922
Haruki Nishimura retweeted
A really solid step toward scalable, high-quality robot data collection — Raiden, from colleagues at TRI @ZakharovSergeyN (and led by @s1wase) lowering the barrier to entry for bimanual data collection, with support for leader–follower setups and SpaceMouse teleop. Big highlight - it natively supports camera calibration and integrates TRI’s learned stereo depth model out of the box, with strong improvements over vanilla ZED SDK. If you're working on robot learning or data collection pipelines, definitely worth a look👇 tri-ml.github.io/raiden/
Our 3D Vision team (3DGR) is releasing Raiden — a data collection toolkit for YAM robots. Built for scalable, high-quality data: supports leader–follower + SpaceMouse teleop, multi-camera setups, and modern stereo depth (incl. TRI learned stereo). tri-ml.github.io/raiden/
2
17
173
15,485