1/N Long horizon, complex tasks that truly matter in everyday life are not solved problems by today’s robotics, requiring planning, object detection, object manipulation, and failure recovery. That's why Stanford's BEHAVIOR Challenge is back for year 2! Last year, the winning solution reached only 12.4% full task success. This year, the BEHAVIOR challenge has more tasks, better evaluation, and is easier to use. 🚨 ⏰ Submission deadline: 10/16/2026 📣 Winners announced: 11/04/2026 🏆 Prize pool: $11,000
22
91
531
115,323
2/N Real-world robot evaluation is essential but hard to scale: experiments are difficult to control, reproduce, and compare. Simulation is a powerful testbed for scalable, controlled, reproducible evaluation. BEHAVIOR-1K is an open-source simulation benchmark of 1,000 everyday household activities requiring long-horizon reasoning, navigation, and bimanual manipulation, giving us a scalable way to measure how well robot AI models generalize. arxiv.org/abs/2403.09227
1
3
45
10,440
💯What’s new in the 2026 BEHAVIOR Challenge? 1. Double the tasks, double the challenge. We doubled the benchmark from 50 to 100 long-horizon household tasks. These activities average 6 minutes each, requiring navigation, planning, memory, and bimanual coordination. No other robotics benchmark comes close in terms of difficulty.
2
3
27
6,164
🔍 2. Larger dataset, better baselines • 20,000 human teleoperation demos, 1950 hours in total (2 times larger than 2025) • RGBD observations and Robot proprioception • Skill/subtask annotations • Strong baseline support: pi0.5, GR00T N1.7
1
1
16
2,817
5/N 🧪 Evaluation & Submission To better reflect real-world deployment, the BEHAVIOR Challenge has one official track this year using only robot onboard observations: • RGB • Depth • Proprioception Submission instructions and evaluation details are available here: behavior.stanford.edu/challe…
1
21
2,462
6/N Together, let’s ask: ❓Can current models solve complete human-centered household tasks? 🔀 How should agents combine control, memory, and planning? 📉 Where do today’s models fail to generalize? 📈 What actually scales in embodied AI?
2
17
2,140
7/N 💬 Join the BEHAVIOR Discord server to ask questions and discuss: discord.gg/bccR5vGFEx We will also hold office hours every Monday, 5–6pm PST over Zoom. See the website for the link. Whether you’re a robotics veteran or just entering the field, we’re here to support you.
1
1
14
1,997
8/N Proud of the amazing work from our students and collaborators, led by @drfeifei: @wensi_ai @stefyfren @cgokmenAI @yalcintur36 @minyeongkim_ @BrndaHere2Chl @AndiXu1111 @RavenHuang4 @RuohanZhang76 @jiajunwu_cs

Jul 13, 2026 · 6:08 PM UTC

1
2
14
1,587
10/N We thank @SimovationInc for providing high-quality JoyLo teleoperation data in simulation for the BEHAVIOR dataset. BEHAVIOR is built upon @nvidia Omniverse. We thank @nvidia for their continuous support.
1
2
30
11,481
11/N We thank our sponsors and supporters for their generous support. @SimovationInc @IMDAsg @StanfordHAI @SchmidtFutures Calder Inc.
1
4
26
11,248
Sort replies: Relevant Recent Liked