GPT-6 Sol 和 Luna 发布,最狠的不是模型变强了,是 OpenAI 终于把 benchmark 换成了「多少钱办成一件事」。
AutomationBench 测的是 agent 跑完一整条 workflow(47 个工具,销售/财务/客服全覆盖):Sol 33.2% 成功率,$0.27 跑完一个任务;Claude Opus 5 是 26.9%,成本是它的 11 倍。
作为天天让 agent 跑长任务的人,这组数字比参数升级实在得多——试错成本降下来,长链条任务才敢放开跑。personal agent 的下半场,拼的可能不是谁更聪明,而是谁跑得起。
不过 vendor benchmark 都要打折看:OpenAI 的对比表里全是旧 Opus 5,而 Anthropic 同一天发布的 Opus 5.5 自称同项 40.0%。两家 CEO 十天前还说要一起"pace the frontier",转头就互相砸场子。这出戏,配得上今天的热搜第一。
你们现在用 agent 跑长任务,感觉最烧钱的是哪一步?
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe.
GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale.
We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.