We tested Opus 5.5, GPT-6-Sol, and GPT-6-Luna thoroughly across 100 open-ended coding and engineering environments.
Our results have significantly diverged from AAII. Some takeaways:
• GPT-6-Astra is still the frontier model by a comfortable margin.
• Opus 5.5 ranges from #2-#7 on coding categories, being the second best model in its ability to one-shot code (which is our best correlated measure with fluid intelligence).
We don't test "usability," but anecdotally, it seems they improved its communication style. And the lower pricing is a welcome change -- we can probably thank
@thsottiaux for that. It's not a cheap model, but it's the first reasonably priced Anthropic model.
However, Opus 5.5 underperformed in our custom harness, which is unusual for Claude models, and a pattern that started with Fable 5.1 (one-shot-fluid intelligence measured better than agentic results). One observation is that it frequently self-stopped when it felt its results were good enough: e.g. in one evaluation, Opus 5.5 closed the session, with the reason: "averaged 867,703 in the latest evaluation, ranked 2nd of 31 against sampled opponents"... so I suppose that saves money, but it still has a tendency to make questionable unilateral judgment calls like that. In our harness, almost every other frontier model keeps working until their code stops improving or they use their full call budget. I do wonder how much of this is due to Anthropic retuning their "High" effort mode (which we use for testing)*.
• GPT-6-Sol is within margin of error of Opus 5.5, and Pareto optimal*. It's faster than all Anthropic models, even on a Flex endpoint. Competitive with the frontier on intelligence and cost.
• GPT-6-Luna is slightly smarter than GPT 5.6 Luna, and discounted. The old Luna already had a lot of applications at its price point, so this model might end up being the most impactful of the group on non-engineering work.
About GBENCH from
@GertLabs: our environments are open-ended with no single correct answers, but verifiable results. They involve creativity, design, and are often organized as multi-agent coding games or engineering design challenges. You can see some live demo examples at
gertlabs.com/spectate. Every one of our environments in the benchmark pool is unsaturated, and we remove any eval that shows any signs of saturation/stagnation across multiple releases. It's designed to measure and differentiate uncontaminated, raw intelligence for frontier models specifically.
These results are interesting and we've thoroughly reviewed them. I don't like to see Opus 5 still so high on the leaderboard, but ultimately I think this is more a problem with Anthropic's newest releases, and it probably explains the price decrease and consecutive releases since Astra came out. They haven't significantly improved upon Opus 5's intelligence, only its personality (which is honestly still a huge win tbh).
*We don't chase marketing terms like xhigh and max. We test all models on adaptive/auto-reasoning where supported and "high" effort where configurable. Some companies beef up their max effort more than others (making them unrealistically slow and expensive in practice), but "high" shows you what an underlying model is capable of and is the more common configuration. This is likely working against Opus 5.5 here.
*Some caveats on price and speed -- we have been using the "OpenAI Flex" endpoint since the Astra release, which I recommend you try if you're using API pricing. It's half price and a little slower, but OpenAI models are already quite fast so it hasn't been an issue in practice. But keep that in mind when making price comparisons to earlier OpenAI models on our cost and speed efficiency charts.