1/N Long horizon, complex tasks that truly matter in everyday life are not solved problems by today’s robotics, requiring planning, object detection, object manipulation, and failure recovery. That's why Stanford's BEHAVIOR Challenge is back for year 2! Last year, the winning solution reached only 12.4% full task success. This year, the BEHAVIOR challenge has more tasks, better evaluation, and is easier to use. 🚨 ⏰ Submission deadline: 10/16/2026 📣 Winners announced: 11/04/2026 🏆 Prize pool: $11,000

Jul 13, 2026 · 6:02 PM UTC

22
91
531
115,311
2/N Real-world robot evaluation is essential but hard to scale: experiments are difficult to control, reproduce, and compare. Simulation is a powerful testbed for scalable, controlled, reproducible evaluation. BEHAVIOR-1K is an open-source simulation benchmark of 1,000 everyday household activities requiring long-horizon reasoning, navigation, and bimanual manipulation, giving us a scalable way to measure how well robot AI models generalize. arxiv.org/abs/2403.09227
1
3
45
10,437
💯What’s new in the 2026 BEHAVIOR Challenge? 1. Double the tasks, double the challenge. We doubled the benchmark from 50 to 100 long-horizon household tasks. These activities average 6 minutes each, requiring navigation, planning, memory, and bimanual coordination. No other robotics benchmark comes close in terms of difficulty.
2
3
27
6,162
🔍 2. Larger dataset, better baselines • 20,000 human teleoperation demos, 1950 hours in total (2 times larger than 2025) • RGBD observations and Robot proprioception • Skill/subtask annotations • Strong baseline support: pi0.5, GR00T N1.7
1
1
16
2,816
5/N 🧪 Evaluation & Submission To better reflect real-world deployment, the BEHAVIOR Challenge has one official track this year using only robot onboard observations: • RGB • Depth • Proprioception Submission instructions and evaluation details are available here: behavior.stanford.edu/challe…
1
21
2,461
6/N Together, let’s ask: ❓Can current models solve complete human-centered household tasks? 🔀 How should agents combine control, memory, and planning? 📉 Where do today’s models fail to generalize? 📈 What actually scales in embodied AI?
2
17
2,140
7/N 💬 Join the BEHAVIOR Discord server to ask questions and discuss: discord.gg/bccR5vGFEx We will also hold office hours every Monday, 5–6pm PST over Zoom. See the website for the link. Whether you’re a robotics veteran or just entering the field, we’re here to support you.
1
1
14
1,997
8/N Proud of the amazing work from our students and collaborators, led by @drfeifei: @wensi_ai @stefyfren @cgokmenAI @yalcintur36 @minyeongkim_ @BrndaHere2Chl @AndiXu1111 @RavenHuang4 @RuohanZhang76 @jiajunwu_cs
1
2
14
1,585
10/N We thank @SimovationInc for providing high-quality JoyLo teleoperation data in simulation for the BEHAVIOR dataset. BEHAVIOR is built upon @nvidia Omniverse. We thank @nvidia for their continuous support.
1
2
30
11,480
11/N We thank our sponsors and supporters for their generous support. @SimovationInc @IMDAsg @StanfordHAI @SchmidtFutures Calder Inc.
1
4
26
11,247
Sort replies: Relevant Recent Liked
Replying to @drfeifei
It just goes to show how incredibly hard it is for a robot to do things we don't even think about like adapting when something goes wrong detecting weirdly shaped objects or just planning a basic household chore
1
699
Replying to @drfeifei
that's exactly where most people go wrong: thinking robots need to solve all those problems at once
470
Replying to @drfeifei
🚨 Top mathematicians just issued a clear warning about AI: Don't believe the hype. Over 2,300 mathematicians, including Fields Medal winners Terence Tao and Peter Scholze, have signed the Leiden Declaration on Artificial Intelligence and Mathematics. Endorsed by the International Mathematical Union, it is the most significant collective response from a major academic discipline evaluating frontier AI impact. The core message is straightforward: current AI tools have real constraints when applied to complex work, and commercial incentives are pushing claims beyond what the technology can reliably deliver. Read the full declaration here: leidendeclaration.ai Why this matters beyond mathematics The declaration identifies five threats that apply to any field deploying AI: 1) Plausible but unreliable outputs. AI produces arguments that "look" correct but contain subtle errors. In high-stakes work, human verification is critical and costly. 2) Attribution collapse. Models trained on published work don't properly cite sources. Training data was often obtained by exploiting licenses or violating copyright protections. 3) Distorted incentives. AI use becomes incentivized for its own sake, warping hiring, funding and recognition. 4) Press release science. Results announced "on market timelines" before community evaluation can take place. Commercial incentives drive firms to "overstate the capabilities of their products." 5) Loss of autonomy. Research priorities shift toward what is automatable rather than what is significant. The leap: chatbots → agentic AI → software → research We have moved from chatbots to agentic systems. Now AI is solving 80-year-old mathematical conjectures. The declaration is not about toy problems. It is about frontier systems being deployed in contexts where correctness matters. What this means for your industry The same risks apply wherever AI is used in high-stakes work: law, medicine, finance, engineering. The declaration's core insight is simple: AI generates narrative, not truth. Verification cannot be automated away. Human accountability is non-negotiable. 💼 I’ve written a more detailed breakdown of how these risks show up in practice and what organisations are doing about them. It’s available for subscribers. ☕️ What have you observed in your industry? Have verification or hidden costs issues already appeared in your AI deployments?
304
Replying to @drfeifei
the 12.4% is the most honest number in robotics right now. everyone posts the highlight reel of the one task that worked; a benchmark that scores full-task completion across 100 long-horizon activities is where the highlight reel goes to die. and moving to onboard-only RGB-D and proprioception is the real tell. no privileged sim state, no overhead oracle. that gap between 12.4% and the demo video is the entire field.
5
601
Replying to @drfeifei
Finally a real work performance benchmark
1
299
Replying to @drfeifei
Is this the ImageNet of robotics?
94
Replying to @drfeifei
李飞飞一直陷在美国式致命自负里!
170
Replying to @drfeifei
The World As I See It
168
Replying to @drfeifei
12.4%. Thank you for publishing that plainly. Same year as IMO gold. That gap isn't a data problem. On 6/N: our bet is that memory can't be a module beside the policy in the brain, the hippocampus sits inside the loop. Six-minute horizons are memory problems wearing a manipulation costume. #SHARM brain-inspired, non-LLM world model. Not a robotics stack, but that's our whole thesis.
98
Replying to @drfeifei
12.4% full-task success is a useful dose of reality. Which failure category dominated last year: long-horizon planning, manipulation precision, or recovery after one bad step?
163
Replying to @drfeifei
👌
162
Replying to @drfeifei
Long-horizon home robotics is the honest benchmark. Planning, manipulation and recovery in one messy task exposes every shortcut.
69
Replying to @drfeifei
Kind of humbling that the winning solution got 12.4%. We talk about AGI constantly but a robot still can't reliably pour a drink in a diner.
60