GPT-6 Luna deserves its own line. At max effort it is the first small hosted model on the PINNACLE board to finish agentic work at a frontier rate, and it does it for 3 cents a correct task.
➡️ 99.6% of the multi-step enterprise jobs finished end to end in our testing, up from 88.6% at medium, the same completion rate as GPT-6 Sol at max
➡️ A Model Score of 4,115, above Claude Sonnet 5 (3,782) and GPT-5.6 Sol (3,954), for 27x and 33x less per correct task, and above every open-weight build on the cost view except Qwen3.8-Flash-Next
➡️ Four rows leave the cost frontier because of it: GPT-6 Sol at medium, DeepSeek-V4-Flash-0731, Qwen3.8-27B and Qwen3.6-27B, so a hosted model now holds every frontier step below 5,000 but one
➡️ The cost of the jump is tokens, 30,624 output tokens per correct answer against 7,332 at medium, and retrieval is still where it shows its size: 92.8% on the 128K corpus against 97.3% for Sol, with fabrication at 11.4%, down from 14.4% but still second highest among the hosted models
OpenAI does not publish a parameter count for Luna, and its $0.10 / $0.50 list price puts it in the small tier. Nothing priced like that has finished agentic jobs at this rate on our board before.