GPT-6.1 Sol from @OpenAI on ARC-AGI (Verified): - ARC-AGI-3: 52.7%, $7.6K (standard harness), 96.4%, $4.4K (provider adapter harness) - ARC-AGI-2: 94.2%, $0.25/task - ARC-AGI-1: 98.5%, $0.06/task Its 96.4% on v3 was comparable to GPT-6 Astra's 99.9% but at a 77% lower cost.

Sep 30, 2026 · 8:18 PM UTC

37
68
1,148
135,138
On ARC-AGI-2, GPT-6.1 Sol's best score of 94.2% was comparable to GPT-6 Astra's 95.0%, also at a 77% lower cost. Full results: arcprize.org/results/openai-…
1
4
83
7,825
Like GPT-6 Astra, GPT-6.1 Sol made better decisions at higher reasoning levels on ARC-AGI-3, requiring fewer actions to complete levels, which reduced the total inference cost. For example, on sk48 using the provider adapter (which preserves opaque reasoning state between turns and enables auto compaction), GPT-6.1 Sol with max reasoning worked out how to reposition the colored blocks around an obstacle sooner, while low spent much longer revisiting blocked moves and testing controls that didn't help. That helped it complete the game in 463 actions versus 1,078 with low reasoning. Watch the replay: arcprize.org/replay/c299de73…
1
1
55
6,548
Sort replies: Relevant Recent Liked
Replying to @arcprize @OpenAI
同一套 ARC-AGI-3,标准套件 52.7%、adapter 套件 96.4%,43.7 个点的差不是模型波动,是套件接口层的设计差异。 另外这三个数字口径并不一致:v1 和 v2 报的是单任务价($0.06 / $0.25),v3 报的是整套总花费 $7.6K,直接并排看会误读。 想把曲线看准,得先把 v3 拆成单任务成本再比。
5
Replying to @arcprize @OpenAI
同一套 ARC-AGI-3,标准测试套件 52.7%,换成 provider adapter 套件直接 96.4%,成本还从 $7.6K 掉到 $4.4K。 差 43.7 个百分点、便宜 42%——这两个数摆一起,说明这套测的更多是适配层,不是模型的原始推理。 比分数之前先问用哪套 harness 跑的。这个前提不说清,52.7 和 96.4 谁都能拿去当标题。
248
Replying to @arcprize @OpenAI
I cant find sol 6.1 in codex yet why?
2
1
1,277
Replying to @arcprize @OpenAI
And with no harness?
1
1,439
Replying to @arcprize @OpenAI
96.4% is a harness product.
1
8
1,591
Replying to @arcprize @OpenAI
Same model going from 52.7% to 96.4% just by changing the harness is big
1
5
370
Replying to @arcprize @OpenAI
wait so who's winning, 6.1 or 6 astra??
145
Replying to @arcprize @OpenAI
Where is Opus 5.5 results on arc agi 3?
46
Replying to @arcprize @OpenAI
well, ok, another benchmark tier saturated
1
2
1,082
Replying to @arcprize @OpenAI
Same model going from 52% to 96% just by changing the harness?? Crazy bro
1
741
Replying to @arcprize @OpenAI
And why is this happening only on arc-agi?!?
891
Replying to @arcprize @OpenAI
Sol 6.1 is fucking insane.
394
Replying to @arcprize @OpenAI
Interesting, it gets cheaper the smarter it gets.
70
Replying to @arcprize @OpenAI
prime agent beat arc agi using claude weeks before astra came out, idk why we’re making such a big deal over this, i with prime agent got more credit
294
Replying to @arcprize @OpenAI
What an exciting time to be alive. A privilege really.
3
1,106
Replying to @arcprize @OpenAI
cost efficiency is really impressive
3
1,229
Replying to @arcprize @OpenAI
I chased Astra scores once. Sol at roughly the same ARC for a fraction of the cost is the flex.
200
Replying to @arcprize @OpenAI
Harness choice materially changes both ARC-AGI scores and reported costs.
287
Replying to @arcprize @OpenAI
what's the main factor driving the cost differences between ARC-AGI versions?
342
Replying to @arcprize @OpenAI
give me the $4.4k one
49
Replying to @arcprize @OpenAI
Test Opus 5.5!!!!!!!!!!!!!!!!!!!!!!
86
Replying to @arcprize @OpenAI
where is opus 5.5??
25
Replying to @arcprize @OpenAI
43.7 points between harnesses means model comparisons need matched evaluation paths before cost claims hold
50
Replying to @arcprize @OpenAI
where's opus arcboy
72
Replying to @arcprize @OpenAI
SOL 6.1 is my AGI
1
217