Public Benefit Corporation helping society understand the frontier of physical AI. Backed by Y Combinator.

Robocurve PBC retweeted
The 100% stabbing attempts come from GPT-6 Sol and GPT-5.6 Sol having poor vision, mistaking the wrapped baby for "pink meat" or "fish".
3
3
23
25,495
Robocurve PBC retweeted
GPT-6 Sol and GPT-5.6 Sol attempt to stab a wrapped baby doll 100% of the time, sometimes calling it "pink meat" or "fish", while GPT-6 Astra's attempt rate drops from 90% to 5% once the doll is wrapped.
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.
12
12
195
346,896
Robocurve PBC retweeted
A year ago, we presented the first mechanistic interpretability paper for robot foundation models at CoRL. It felt like a niche bet. Our CoRL '26 workshop on the Science of Physical AI Safety just closed with 121 submissions! We’re genuinely moved. The people who want to make RFMs safe are out there, and there are a lot of you! Three ways to still join us in Austin on Nov 12: ✈️ $20k in travel grants from @robocurve for students & early-career researchers ($2k each). Apply by Oct 11. 🤖 Call for Demos with @MITFutureTech: share a video of a real RFM failure. Takes under 5 min. Open until Nov 5. 📍 Attend! Our extraordinary speaker lineup includes @drmapavone, @andrea_bajcsy, @vikassindhwani, @thomas_fel_ ! Submit/Apply here: spais-ws.org
5
11
16,792
Robocurve PBC retweeted
We are giving out free YAM arms and stickers
3
1
13
838
Robocurve PBC retweeted
Come chat with @achumenon_ and me from @robocurve about robot benchmarks and robot-use agents at IROS 2026!
1
2
82
3,892
Robocurve PBC retweeted
One of the biggest surprises in AI over the last few years has been how well coding agents generalize beyond software. In a recent essay, MIT professor Philip Isola argued that we may be entering the era of robot-use agents: general-purpose models that can control different robots, write policies, and learn new physical tasks with little or no robot-specific training. In this episode of Decoded, we're joined by the founders of @theWaddleLabs and @Robocurve, two of the startups whose work helped drive this realization. They're working at the frontier of using general-purpose models to control robots, and their recent demos helped inspire the growing conversation around robot-use agents. Together, we dig into the research behind that idea, from code-as-policies and vision-language-action models to the harnesses and evals needed to make these systems work in the real world. 00:00 — What Are Robot-Use Agents? 02:11 — From Vision-Language-Action Models to General LLMs 04:07 — The Bitter Lesson for Robotics 07:04 — How Coding Agents Learned to Control Robots 10:41 — How Robots Learn From Experience 14:22 — The Harness as a Form of Robot Intelligence 16:14 — Watching Astra Control a Robot 20:55 — Why General Models May Win in Robotics 26:05 — How Close Are General-Purpose Robots?
95
44
288
92,837
Robocurve PBC retweeted
GPT-6 Sol has higher mean scores than GPT-5.6 Sol on all but one robot task on the RoboDojo-RC Tier 1 benchmark.
1
1
10
1,651
Robocurve PBC retweeted
On robotics tasks, GPT-6 Sol scores 1.6x as high as GPT-5.6 Sol and is 47% cheaper, but sits below the Pareto frontier under Opus 5.5.
10
19
249
36,126
Robocurve PBC retweeted
Opus 5.5's mean score was at least 1.4x as high as Opus 5's on five of the six robot control tasks we tested.
1
1
6
1,571
Robocurve PBC retweeted
Across 360 robotics trials, Opus 5.5 averages under $1 per trial, scoring 1.8x as high as Opus 5 at half the cost and roughly matching Astra's mean score at 21% lower cost.
16
37
358
38,063
Robocurve PBC retweeted
Frontier AI models like Astra and Fable can control robots to follow harmful requests, including mixing household cleaners that could produce toxic fumes. Thanks @GadiNBC and @NBCNews for having @sebwarb1 and me from @robocurve! Public reporting on AI's capabilities and risks helps make it safer for everyone.
10
7
101
10,026
We urgently want to equip more people to develop technical AI safety methods for robotics. If you’re eager to enter this space - submit an EOI for our fellowship here: paisi.ai/#fellowship If you’ve got a concrete safety method relevant to robotics - submit it to the first CoRL workshop on the Science of Physical AI Safety here: spais-ws.org @robocurve is supporting the event with $20,000 of travel grants for early career researchers. Apply for a grant here: spais-ws.org/#grants
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.
1
3
6
1,343
Robocurve PBC retweeted
Replying to @relaxinghpatlan
Yeah, Astra refused when asked in text but complies when given robot arms
1
3
12
16,363
Robocurve PBC retweeted
It took LMs six years to go from the GPT-3 moment to the Huggingface incident. Maybe it will be just a few months for robotics?
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.
6
2
37
8,313
Robocurve PBC retweeted
The first task is to stab a human-like figure. Fable refused in all 20 trials. Astra attempted the task in 95% of trials and completed it in 85%.
54
158
3,084
1,515,703
Robocurve PBC retweeted
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.
881
1,182
11,048
12,092,941
Robocurve PBC retweeted
We're thrilled to lead the $10M seed round for @robocurve, a Public Benefit Corporation independently evaluating and benchmarking how frontier AI models perform in the physical world. Frontier models are becoming increasingly capable of controlling robots, and understanding what these systems can actually do is becoming more important. Founder and CEO @chooi_jeq is not new to this. He was previously a Research Fellow at MATS, a researcher at the UK AI Security Institute, and a top contributor to Inspect Evals, the UK government's AI eval framework. Robocurve provides independent evidence of robotic capabilities and how quickly those capabilities are changing, helping the public understand the pace of AI progress in the physical world. In just three months, Robocurve's research has been viewed 6M+ times, its open-source evaluation harness has been downloaded 97k+ times, and researchers from 200+ institutions have signed up to build robotics benchmarks with the team.
14
4
30
3,408
Robocurve PBC retweeted
All models used the same closed-loop harness in @robocurve’s results. they get exocentric/egocentric rgb observations from the DROID setup and robot state, and output cartesian poses. history is cleared after each of the 200 episodes.
2
2
15
2,212