1/N Long horizon, complex tasks that truly matter in everyday life are not solved problems by today’s robotics, requiring planning, object detection, object manipulation, and failure recovery.
That's why Stanford's BEHAVIOR Challenge is back for year 2! Last year, the winning solution reached only 12.4% full task success. This year, the BEHAVIOR challenge has more tasks, better evaluation, and is easier to use. 🚨
⏰ Submission deadline: 10/16/2026
📣 Winners announced: 11/04/2026
🏆 Prize pool: $11,000
22
91
531
115,313
2/N Real-world robot evaluation is essential but hard to scale: experiments are difficult to control, reproduce, and compare. Simulation is a powerful testbed for scalable, controlled, reproducible evaluation.
BEHAVIOR-1K is an open-source simulation benchmark of 1,000 everyday household activities requiring long-horizon reasoning, navigation, and bimanual manipulation, giving us a scalable way to measure how well robot AI models generalize.
arxiv.org/abs/2403.09227
Jul 13, 2026 · 6:03 PM UTC
1
3
45
10,438
💯What’s new in the 2026 BEHAVIOR Challenge?
1. Double the tasks, double the challenge.
We doubled the benchmark from 50 to 100 long-horizon household tasks. These activities average 6 minutes each, requiring navigation, planning, memory, and bimanual coordination. No other robotics benchmark comes close in terms of difficulty.
2
3
27
6,162
🔍 2. Larger dataset, better baselines
• 20,000 human teleoperation demos, 1950 hours in total (2 times larger than 2025)
• RGBD observations and Robot proprioception
• Skill/subtask annotations
• Strong baseline support: pi0.5, GR00T N1.7
1
1
16
2,816
5/N 🧪 Evaluation & Submission
To better reflect real-world deployment, the BEHAVIOR Challenge has one official track this year using only robot onboard observations:
• RGB
• Depth
• Proprioception
Submission instructions and evaluation details are available here: behavior.stanford.edu/challe…
1
21
2,461
6/N Together, let’s ask:
❓Can current models solve complete human-centered household tasks?
🔀 How should agents combine control, memory, and planning?
📉 Where do today’s models fail to generalize?
📈 What actually scales in embodied AI?
2
17
2,140
7/N 💬 Join the BEHAVIOR Discord server to ask questions and discuss:
discord.gg/bccR5vGFEx
We will also hold office hours every Monday, 5–6pm PST over Zoom. See the website for the link.
Whether you’re a robotics veteran or just entering the field, we’re here to support you.
1
1
14
1,997
8/N Proud of the amazing work from our students and collaborators, led by @drfeifei:
@wensi_ai
@stefyfren
@cgokmenAI
@yalcintur36
@minyeongkim_
@BrndaHere2Chl
@AndiXu1111
@RavenHuang4
@RuohanZhang76
@jiajunwu_cs
1
2
14
1,586
9/N And with the strong support from
@EvansXuHan
@jin_lynn808
@Hang_Yin_
@ChengshuEricLi
@josiah_is_wong
@sanjana__z
@YunfanJiang
@wenlong_huang
@RobobertoMM
@YunzhuLiYZ
@ManlingLi_
@Weiyu_Liu_
@silviocinguetta
@hyogweon
Prof. Karen Liu
2
1
12
1,580
10/N We thank @SimovationInc for providing high-quality JoyLo teleoperation data in simulation for the BEHAVIOR dataset.
BEHAVIOR is built upon @nvidia Omniverse. We thank @nvidia for their continuous support.
1
2
30
11,480
11/N We thank our sponsors and supporters for their generous support.
@SimovationInc
@IMDAsg
@StanfordHAI
@SchmidtFutures
Calder Inc.
1
4
26
11,247


