GPT-6 Sol and Luna just landed in Astra’s orbit. Both launch today with API prices 50% lower than GPT-5.6. Build with Sol. Scale with Luna. To production and beyond.
283
754
9,652
980,412
Improvements in caching and inference let us deliver GPT-6 Sol and GPT-6 Luna at lower cost. We’re passing those savings on through lower API prices. Sol and Luna deliver more intelligence at half of the previous generation’s API prices, making more complex work practical to automate at scale.
13
56
783
80,029
On DeepSWE v1.1, which tests coding agents on long-horizon engineering tasks, GPT-6 Sol (max) nearly matches Claude Fable 5 (xhigh) at ~80% lower cost per task. GPT-6 Luna (max) is comparable to Fable 5 (medium) at 96% lower cost per task.
19
33
555
78,676
On AutomationBench, which measures whether AI models can complete real business workflows across apps, GPT-6 Sol (xhigh) leads Fable 5.1 (max with Opus 5 fallback) with 88% lower reported cost per task.

Sep 22, 2026 · 6:14 PM UTC

6
5
175
22,445
With improved prompt caching in GPT-6, agents do less redundant processing behind the scenes. More context stays reusable as reasoning effort and tool availability change, making responses faster and cheaper.
4
4
178
60,659
GPT-6 Sol and GPT-6 Luna build on Astra’s alignment advances, showing improvements over the GPT-5.6 Sol and GPT-5.6 Luna models in our alignment evaluations.
Replying to @OpenAI
GPT‑6 Sol and Luna build on Astra’s advances in alignment, showing improvements over their GPT-5.6 counterparts.
4
4
174
37,930
Sort replies: Relevant Recent Liked
Replying to @OpenAIDevs
88% lower cost per task changes what teams will actually automate. Workflows that were too expensive to run continuously can start becoming normal infrastructure.
119
Replying to @OpenAIDevs
Que un agente complete flujos empresariales reales con 88% menos costo es relevante para la contaduría y las finanzas. La adopción responsable exige evidencia de cada paso, trazabilidad de las fuentes y controles de revisión humana antes de afectar registros, cierres o decisiones de auditoría.
51
Replying to @OpenAIDevs
Someone called this the worst benchmark graph ever, and I agree.
27
Replying to @OpenAIDevs
For network ops agents, one failed DNS change can erase a cheap token bill. I'd track cost per accepted workflow, retries, tool failures and rollback rate. This benchmark result is promising; production needs the same acceptance criteria and failure budget.
13