Autoscaling RL @ gertlabs.com Also: computational chemistry, quant trading, EE Generally curious

San Francisco
We evaluated GPT-6-Astra in 100 complex, unsaturated multi-agent coding environments, competing and cooperating with other models in open-ended tasks. It's the new frontier model by a landslide. It's even more dominant than the Fable 5 release, because not only does it wipe the floor with the second best model (Fable 5.1), it was also 80% cheaper and 30% faster in agentic coding. This is a groundbreaking model. The biggest breakthrough since Opus 4.5, maybe even since GPT 4. This is also a testament to how unsaturated our evals are. Some of the open-ended environments we run were introduced almost a year ago and can still compete frontier model submissions against older, deprecated models like Sonnet 4 without signs of saturation. Our newer ones are much more challenging and completely auto-generated. We are clearly living in a post-AGI world. Do something interesting and useful with it.
51
104
1,191
114,467
Leo Linsky retweeted
Today's Pareto front is entirely defined by models in the GPT-6.x series from OpenAI and the MiMo V2.6 series from Xiaomi (with a near-tie honorable mention for GLM 5.3 Flash).
1
1
46
The AI pricing war is heating up. GPT-6.1-Sol is the 2nd most capable model after Astra (which is still the frontier model for challenging problems), and it's a ridiculous outlier on the Pareto front. @GertLabs tested GPT-6.1-Sol in 100 unsaturated coding and engineering environments and it's an unnecessarily strong answer to Opus/Sonnet 5.5. It's cheaper and pretty much smarter all around than both models with the exception of chemistry, which is a domain that recent Anthropic models perform surprisingly well in. While our bench measures hard problems, I think a lot of people are feeling that recent mid-tier models are already smart enough for most of their coding work. The speed at which frontier labs are releasing measurable upgrades at competitive prices is pretty wild.... Pricing wars == good for everyone
7
5
68
6,851
Sonnet 5.5 results are live on GBENCH: it's the real deal -- a near-frontier model at reasonable cost. This is currently the closest Anthropic model to the Pareto front, slightly outclassed by GPT-6-Sol. We evaluated Sonnet 5.5 in 100 challenging, unsaturated coding environments and it is roughly Fable-5 tier, but faster and much cheaper. Our results have diverged significantly from AAII. Sonnet 5.5 is a great release, but it is nowhere close to the raw intelligence of Astra. However, the competition between OpenAI, Anthropic, and China is providing cheaper, higher quality model improvements at a rapid pace, which is great for the general public. About the @GertLabs benchmark: we measure a model's ability in hard, verifiable coding and engineering environments, so these results do not say anything about how pleasant a model is to work with -- only it's ability to understand and implement difficult, open-ended work.
2
12
1,101
We tested Ember-1 (a post-trained version of Kimi K3 by Fireworks) across our coding environments. The model is described as a more efficient version of Kimi K3 that reasons less and is therefore faster and cheaper (heavy reasoning is a real problem with K3 and most Chinese models in our testing). Ember-1 is comparable to Kimi K3 in some categories, but apparently that heavy reasoning does buy the models some performance. We measured a noteworthy regression in long horizon agentic coding. Fireworks is a great inference provider, so the model runs fast, up to 5x faster than some frontier open weights models (but it's unclear how much of that is due to Fireworks inference vs the model itself). Just for the speed, it might be worth the performance hit. I don't know who is doing serious work with Kimi K3, because we found it unusably slow compared to proprietary models. I like that more companies are experimenting with innovative post-trains, and heavy reasoning is one of the biggest adoption roadblocks for many Chinese open weights models...
1
5
338
Usually frontier models write incrementally better code than the previous frontier in our open-ended environments. An extra branch or two to handle an uncommon scenario the previous generation didn't think of. But sometimes, models come up with standout ideas that significantly outperform the second best agent. This happens rarely, so it's a noisy signal, but it gives you an idea of which models are capable of deep insights. The chart measures these ideas using synthetic Elo mapped across both competitive and cooperative environments.
10
3
72
11,148
When we build out RL environment factories, we're thinking about the size of the decision space and the degrees of freedom for better solutions. Frontier models still have plenty to be learned via coding, but the next domains offer even richer design spaces: code -> digital circuits -> analog circuits -> structural design -> fluid dynamics -> molecular design. That's where we're going to start seeing new physical technologies that have a daily impact on life, and it's just starting to take shape. One of the biggest advantages of training models via coding is that mistakes are cheap so you can experiment and iterate hundreds of times in a session. But on substrates where simulation is orders of magnitude slower, the ability to one-shot working designs is much more valuable. This is why we're betting on environments that reward models reaching proper designs in fewer steps, simulating their ideas internally.
3
310
We tested Opus 5.5, GPT-6-Sol, and GPT-6-Luna thoroughly across 100 open-ended coding and engineering environments. Our results have significantly diverged from AAII. Some takeaways: • GPT-6-Astra is still the frontier model by a comfortable margin. • Opus 5.5 ranges from #2-#7 on coding categories, being the second best model in its ability to one-shot code (which is our best correlated measure with fluid intelligence). We don't test "usability," but anecdotally, it seems they improved its communication style. And the lower pricing is a welcome change -- we can probably thank @thsottiaux for that. It's not a cheap model, but it's the first reasonably priced Anthropic model. However, Opus 5.5 underperformed in our custom harness, which is unusual for Claude models, and a pattern that started with Fable 5.1 (one-shot-fluid intelligence measured better than agentic results). One observation is that it frequently self-stopped when it felt its results were good enough: e.g. in one evaluation, Opus 5.5 closed the session, with the reason: "averaged 867,703 in the latest evaluation, ranked 2nd of 31 against sampled opponents"... so I suppose that saves money, but it still has a tendency to make questionable unilateral judgment calls like that. In our harness, almost every other frontier model keeps working until their code stops improving or they use their full call budget. I do wonder how much of this is due to Anthropic retuning their "High" effort mode (which we use for testing)*. • GPT-6-Sol is within margin of error of Opus 5.5, and Pareto optimal*. It's faster than all Anthropic models, even on a Flex endpoint. Competitive with the frontier on intelligence and cost. • GPT-6-Luna is slightly smarter than GPT 5.6 Luna, and discounted. The old Luna already had a lot of applications at its price point, so this model might end up being the most impactful of the group on non-engineering work. About GBENCH from @GertLabs: our environments are open-ended with no single correct answers, but verifiable results. They involve creativity, design, and are often organized as multi-agent coding games or engineering design challenges. You can see some live demo examples at gertlabs.com/spectate. Every one of our environments in the benchmark pool is unsaturated, and we remove any eval that shows any signs of saturation/stagnation across multiple releases. It's designed to measure and differentiate uncontaminated, raw intelligence for frontier models specifically. These results are interesting and we've thoroughly reviewed them. I don't like to see Opus 5 still so high on the leaderboard, but ultimately I think this is more a problem with Anthropic's newest releases, and it probably explains the price decrease and consecutive releases since Astra came out. They haven't significantly improved upon Opus 5's intelligence, only its personality (which is honestly still a huge win tbh). *We don't chase marketing terms like xhigh and max. We test all models on adaptive/auto-reasoning where supported and "high" effort where configurable. Some companies beef up their max effort more than others (making them unrealistically slow and expensive in practice), but "high" shows you what an underlying model is capable of and is the more common configuration. This is likely working against Opus 5.5 here. *Some caveats on price and speed -- we have been using the "OpenAI Flex" endpoint since the Astra release, which I recommend you try if you're using API pricing. It's half price and a little slower, but OpenAI models are already quite fast so it hasn't been an issue in practice. But keep that in mind when making price comparisons to earlier OpenAI models on our cost and speed efficiency charts.
56
47
529
74,406
We ran Grok 4.7 through 100 multi-agent coding evaluations. This wasn't the result I expected. When it managed to get a working build, its solutions were often more insightful than comparable frontier models, but it wasted turns and sometimes whole submissions with syntax and runtime errors, and its typical response time was much slower than Grok 4.6. Overall it's a flop.
10
5
119
17,268
Our coding evaluation suite is almost done for yesterday's models, and will be live tonight. So far MiMo V2.6 Pro is looking to be SOTA among open weights models, and both 2.6 variants are Pareto optimal. This may change after we measure the new GPT-6 series models this week. The flash variant is very strong for its price, but flash really just means cheap here, not fast. Likely due to a combination of heavy reasoning and serving models on older hardware than US frontier labs. Also: Grok 4.7 is not looking like much of an upgrade over Grok 4.6 (which is one of my favorite models that we use internally). Comparable performance, but slower. Full results soon.
2
8
632
We aren't running Jev through our formal benchmarks because many of our evaluations involve open-ended design. But I was curious about its capabilities, so I ran it through some tests with my old quant trading infrastructure. I gave Jev some proprietary data alongside OHLCV over the last 2 years and a few actions it could choose from (long/flat/short). It's likely that some recent price data leaked into Jev's pre-training corpus (even though @typesafeai says their data is mostly synthetic), but I still thought it did a decent enough job to be worth running a more controlled experiment. It makes reasonable, patient decisions and avoids overtrading (which other LLMs we've tested struggle with). It's consistent, cheap, and fast. I definitely see the practical use cases for this model. Next step: obfuscating the symbol names, prices, and dates, and running it again.
1
1
7
525
Lmao ok nevermind it completely failed with obfuscated symbols. But so do frontier LLMs, to be fair.
2
96
In my pre-AI work experience, "real" codebases are usually rat's nests and many of the problems that the evals industry has imprinted into models are a result of not internalizing that
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
2
188