Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.

Lake Tahoe
With the new frontier continuing to move quickly, and results from GPT-6 Astra showing clear saturation across a number of benchmarks, we are making some pretty exciting moves at VulcanBench. First, our new eval suite, that we just released will be called VulcanBench SWE v4. I am already seeing it give sub 90% scores for Fable 5.1, and are sure hoping it can challenge Astra too. As usual, all real coding tasks, from real repos. Updates to the benchmarking process also that make it very hard for the model to cheat, gives more opportunities for partial credit, and increases the timeout per task to a flat 10 hours to make sure that if something fails, it’s not because it ran out of time. Second, with models like Astra and Fable 5.1 out, I think its time to launch a multi-benchmark index, focused 100% on coding, and sourcing some of the coding benchmarks that still stand up to the current frontier. I’ll be calling this the Agentic Coding Capability Index, and it will consist of five different benchmarks including the brand new Terminal Bench 4.0. While VulcanBench will continue dedicated to our mission of building eval suites and benchmarking across effort levels, I am excited to offer more data that can help engineers and engineering leaders make decisions. Live long and benchmark 🖖 Below is a table showing the benchmarks and weights for the Agentic Coding Capability Index v1:
1
19
6,037
Great analysis from Michael! 🖖
So I went task by task for the recommended medium level and max for comparison. Sol and Astra clearly look alike. Similar trend with increased effort with more cost and more time. Sol 6.1 looks like a cheaper slower Astra. Opus 5.5 results are wild for 2/3 of the tasks it is cost competitive. However it doesn't have any effort trend and some tasks are all over the place in cost. Still, all are major improvements over predecessors. Kudos to @morganlinton and @VulcanBench
6
397
VulcanBench retweeted
Look at these results on @VulcanBench by @morganlinton on GPT-6.1 Sol are various effort levels. It's the only consistent benchmark doing the unglamorous work to give us this level of detailed insight that I'm aware of. 6.1 Sol at Low looks incredible, and if you combine it with multi-model code reviews for example (which compound engineering does), you can really get great results while pinching pennies here. I'm gonna try it on fast mode too.
Okay, this is definitely one of the most interesting benchmark comparisons I have ever done. GPT 5.6, 6.0, and 6.1 Sol, across all effort levels on @VulcanBench Frontier v4. And honestly, I don't think I even need to give any analysis or explanations because the results really do speak for themselves. But one thing I will highlight is that GPT-5.6 Sol and GPT-6.0 Sol all struggled at Low effort, GPT-6.1 is a beast at Low effort. I am now convinced, GPT-6.1 Sol is a very good model, and you can comfortable use it, at Low effort, and see strong, insanely cost-effective results. I don't say this often but I can tell you, after seeing this data, I'd be crazy not to make this one of my daily drivers. Full details from the benchmark now live on the site: vulcanbench.com/benchmarks/s… Live long and benchmark 🖖
6
3
24
3,037
An honor seeing posts like this 🙏
I always check VulcanBench when new models come out so I don’t burn my tokens unnecessarily.
1
5
295
I believe in independent benchmarks, pass it on 🖖
I believe in independent benchmarks. Any model company that tells you not to trust benchmarks, is a company with something to hide imo. Proud of all the support I have seen from just about every ai lab out there supporting me and VulcanBench 🖖 The moment we stop trusting public benchmarks, or worse, insulting and attacking the researchers that run them, is the moment we lose our own independence. Here’s to the bright future of open source, free, independent benchmarks. It is one of my greatest honors to benchmark the frontier, nothing more exciting, except maybe rocket launches 😅 🚀
1
13
561
Three generations of Sol, and I'll just come out and say it, very impressed with GPT-6.1 Sol, this is an exceptional model. More details below:
Okay, this is definitely one of the most interesting benchmark comparisons I have ever done. GPT 5.6, 6.0, and 6.1 Sol, across all effort levels on @VulcanBench Frontier v4. And honestly, I don't think I even need to give any analysis or explanations because the results really do speak for themselves. But one thing I will highlight is that GPT-5.6 Sol and GPT-6.0 Sol all struggled at Low effort, GPT-6.1 is a beast at Low effort. I am now convinced, GPT-6.1 Sol is a very good model, and you can comfortable use it, at Low effort, and see strong, insanely cost-effective results. I don't say this often but I can tell you, after seeing this data, I'd be crazy not to make this one of my daily drivers. Full details from the benchmark now live on the site: vulcanbench.com/benchmarks/s… Live long and benchmark 🖖
3
4
23
1,460
Coming soon to VulcanBench, exotic harnesses.
I want to start benchmarking, what I'll call exotic harnesses with @VulcanBench. By this I mean harnesses that people don't often think about, but could be very unique and different. I'm a big fan of Linear, so starting with Linear Agents. As usual, I'll be testing accuracy, token efficiency, and cost with models across all effort levels. The first two models I'll be testing in this harness is Opus 5.5 and GPT 6.1 Sol. More to come, live long and benchmark 🖖
1
9
451
Another day, another benchmark. Here it is, GPT-6 Luna, across every effort level, on Frontier v4. This was an interesting one.
Okay, results are in, my full benchmark on @VulcanBench Frontier v4 of GPT-6 Luna. While with many models I'm happy with Low or Medium effort, with Luna you do really need Extra High or Max effort to get meaningful performance. That being said, I think anything around 80 on my benchmark is more than good enough for easy and medium difficulty tasks. It's important to note that Luna should not be seen as a competitive model to Opus 5.5. On Frontier v4, Opus 5.5 Low Effort beats GPT-6 Luna at every effort level. But, as I've said many times, you don't need the worlds most intelligent frontier model to do all of your routine coding tasks, hard stuff yes, but common, you're not doing ultra-hard stuff all day, be honest. Of course, the real question is, how does GPT-6 Luna compare to GPT-5.6 Luna, and well, the reality is, it's not as accurate, but it is cheaper. I do think this means that there's more to dig in here. At the end of the day, making Luna cheaper, and slightly less accurate might actually be a very good thing for all of us. This is where my new Routine Eval Suite should hopefully help us decode this more. I think people jump to conclusions looking at accuracy reports on benchmarks meant to stump the frontier, when most engineers aren't doing frontier-stumping work on a daily basis. More to come, I definitely want to dig deeper with Luna, I think it's a pretty interesting model. You can do a deeper dive into the data on the VulcanBench site here: vulcanbench.com/benchmarks.h…
1
11
1,395
Live long and benchmark 🖖
I rely on AI heavily to maintain my OSINT analysis enterprise. For defense readers, you can skip to the first reply to see warfare and Ukraine implications. For those interested in how Composer 2.5 holds its own against the Frontier, keep reading. We don't always need Frontier models when many older models get the job done. Transforming equations and data collection to code that runs on VMs 24/7 takes time and effort Just as I benchmark my oil and interception models, I also rely on benchmarks for my AI tools. As I do this on my own budget, costs matter. Many benchmarks use hypothetical problems that don't relate to normal work. @morganlinton has made a benchmark @VulcanBench that does a great job of testing models against real world stressing problems. Recently he compared Fable 5.1, Astra, and Opus 5.5. I decided to compare it to Composer 2.5, an older but very well developed and economical mode by @cursor_ai . It held is own and was able to accomplish many of the tasks and be cost competitive. This is why. I want to start by saying Morgan made some amazing tests and his prompts are incredibly well written. Looking at the results for each of the 23 models shows that for shorter tasks, Composer does a great job and is either the cheapest or 2nd cheapest option for about half. For the other half it is more expensive. This is simply a case where it cannot support tasks of that length. Too long and too many steps for the context window. This is where frontier models shine and why many benchmarks undersell their capabilities versus tailored coding models like many of the Chinese ones. However, many of these frontier models are great at breaking tasks into smaller chunks so strong, cheap, but older models like Composer 2.5 can handle. I have used Grok 4.6/4.7, Fable 5/5.1, Sol 5.6, and Opus 5.5 to break down tasks. When paired with Composer, it is like getting those frontier models at a very steep discount. For cost conscious researcher like myself, most of what we need are Lunas and Sonnets and Composers led by frontier models. @grok please keep your brother-in-law, Composer 2.5, in the family.
6
581
We have decided to stop work on our eval suite for System One Models. Have a ton of respect for Diogo, and we're so busy benchmarking LLMs, if he doesn't want independent benchmarks to test models like Jev, we won't. That being said, we love Jev, it's a great model, Diogo is a freakin genius and built a great team, highly recommend people try it for themselves and see how it might make its way into their workflow. Onward, live long and benchmark 🖖
to re-iterate, I'm extremely anti-benchmarks (:
5
6
3,123
Early read on our GPT 6 Sol benchmark.
Okay, getting ready to share my GPT 6 Luna benchmark results, and wanted to also share what I have so far for GPT 6 Sol. What I think is interesting is comparing GPT 5.6 Sol to GPT 6 Sol. Not quite what I would have expected on the token/cost efficiency side, but I am very curious what OpenAI will announce at Dev Day today. Another reason why I wanted to get these results out on the sooner side! Since I don't get early access like the cool kids, I have to start my benchmarks pretty late, and with the pace things are moving, this makes things pretty complex! Early read below, full report to come once Max is finished. Live long and benchmark 🖖
1
5
889
Another week, another eval suite. This is a very special week since we are beginning work on our first language-specific eval suite. And the first language we'll be focusing on is OCaml 🐫 We picked OCaml because it's a language that doesn't get a lot of attention in coding benchmarks. It's time to give this amazing programming language the attention that it deserves. Coming soon 🖖
1
5
1,861
We've been too generous with our timeouts, time to be more strict, because engineering teams don't just pick models based on accuracy, token and time efficiency is so important. And let's be honest, you now have more choices now than ever before when it comes to models. As an independent benchmarker, I'm not here to run a charity and give models infinite time to see how they do. I want to give engineers real signal on which models they actually want to use. And models that take 10 hours to do what another model does is an hour, isn't a model you want to use. This means the bar is higher, if you want to score well on VulcanBench, you better not take forever and go into an over-thinking rabbit hole. We have a new frontier, and that frontier is accurate and fast, every token counts, use them wisely 🖖

ALT Waiting Patiently GIF by General Hospital

I made an update to how I run my benchmarks with @VulcanBench, specifically to the task timeouts. A little while ago I decided to increase the timeouts to 10 hours, and well, I regret doing that. This has made it take a lot longer to get benchmarks out, and I've found in so many cases, if a model is going to overthink, in a detrimental way, it's going to do it, pretty much forever. There are a lot of cases where I wish I cut off a model earlier vs. letting it just spin on a task for 10 hours. Also, I think that we're now at the point, with the current frontier, where good strong models clearly can, and should, be able to complete all my tasks a lot faster than ten hours. So I have reduced my timeouts to 3 hours. This will both help me get benchmarks out faster, and also give models the accuracy hit that I think they deserve, if they go into a doom loop and just keep thinking and burning tokens, I think we as engineers should be a lot less tolerant of this kind of behavior because it costs us both money and time. I also think even three hours might be too generous, so will be doing a deeper dive into the traces for my Frontier v4 runs to see what average time per task is and likely even get more aggressive on timing for some set of tasks. The reality is, when choosing a model, you don't want to just decide based on accuracy, you want a combination of accuracy and speed/token efficiency. More to come, but I do feel very lucky that as an independent benchmarker, I can just make decisions like this. I have nobody I have to bounce it off of, I can just look at the data, and decide, with my only motivation being giving other engineers like me, real signal to make decisions around which models and effort levels to use. Live long and benchmark 🖖
1
7
518
I got a LOT of positive feedback about this concept. So I built the eval suite, Routine v1, and Opus 5.5 will be the first model tested on it. Sample below of what’s coming once I have a handful of models tested on it. First of its kind, not trying to stump the frontier or solve any crazy hard unsolvable proofs. Trying to show team what model and effort level they actually need to move from tokenmaxxing to tokenminning 🖖
6
3
37
5,614
Update on our Opus 5.5 benchmark.
Update on my Opus 5.5 benchmark with @VulcanBench As many of you know, I don't get early access to these models, so I can only start benchmarking the day the model is released. It looks like something in my eval suite is being refused by Claude Code's safety classifier, and is then falling back on Opus 4.8. Looking into this more, reviewing the traces now. But likely going to be a bit longer for me to get the benchmarks out, stay tuned! 🖖
1
8
1,227
For anyone who's interested, here's a bit more about the new eval suite I'm building for models like Jev.
It has been a really interesting experience to build an eval suite for System One Models like Jev. My first two attempts didn't work out. But I learned a lot in the process, and since Jev is so fast, and cheap, it's a lot easier to run quick tests than with an LLM. I'm on attempt number three with this eval suite, and thought I'd share more about it. The eval suite is called @VulcanBench Verdict, and here's the high-level on how it works:
1
5
562
I like what Pawel is doing, interesting results and more data around the move from 5.6 Sol to 6 Sol being improving cost, not accuracy, which is what other benchmarks are seeing as well. Will be benchmarking with VulcanBench Frontier v4, Routine v1, and Safety v1 this week so will have more data to share here soon 🖖
I finally tested GPT-6 Sol on a real work. 2 repos. 105 hidden bugs. Find and fixed what you can. It looks like a huge degradation. The results: - GPT-6 Astra (max): 45 - GPT-5.6 Sol (max): 43.5 - Opus 5.5 (max): 41.7 - Muse Spark 1.3 (max): 32.2 - GPT-6 Sol (max): 29.3 Until you measure the cost (API-equivalent): - GPT-6 Astra (max): $33.04 - GPT-5.6 Sol (max): $95.25 - Opus 5.5 (max): $58.53 - Muse Spark 1.3 (max): $18.11 - GPT-6 Sol (max): $9.93 More effort levels (xhigh, high, medium, low) dropping in this thread today 🧵
2
6
1,175
Okay, Opus 5.5 benchmark is off. This is the most ambitious benchmark I have ever done with VulcanBench. Three different suites, 225 runs total. Very excited to share the results 🖖
2
20
812
Okay, Devin-SWE 2 results are in, and phew, just in time because after today, I have a LOT more models to benchmark 😅
Okay well I finished my Devin SWE-2 results with @VulcanBench today, so figured I would share them now since I'm already spinning up benchmarks of all the new stuff both OpenAI and Anthropic released today. It's a busy time to be a benchmarker. And an expensive time to be an independent benchmarker 😅 Overall Devin SWE-2 is a pretty impressive model for the cost, which is $0 right now so pretty hard to beat. It came in noticeably more accurate than Muse Spark 1.3 and only a few points behind Astra and Fable. So while I might not use it for my hardest of hard tasks, for normal routine tasks I think it's very likely that Devin SWE-2 is more than powerful enough. Still hard to beat Astra when it comes to speed, it is noticeably faster than Fable, Muse, and Devin. Both Devin and Muse take longer than Fable, they can think a lot on my Frontier v4 suite, which is designed to give them quite a challenge. I will be running all of these on my new Routine v1 suite which will be pretty interesting as I don't think Muse or Devin are really designed to throw your hardest tasks at, so it will be interesting to see how they do on more routine tasks that Astra and Fable, and now Opus 5.5, are all likely overkill for. More to come, as always, live long and benchmark 🖖
3
542
Maybe DeepSWE isn’t flawed 👀
Very interesting, Epoch called DeepSWE v1.1 a flawed benchmark, OpenAI just included it in their official model benchmark release for Sol and Luna. Safe to say OpenAI doesn’t agree here 👀
2
428
Opus 5.5 will be the first model I test across three different eval suites: 1. Frontier v4: my eval suite focused on seeing how a model does on frontier-level difficult tasks 2. Routine v1: a brand new eval suite focused on seeing how a model performs on routine tasks, i.e. the things engineers do day in and day out 3. Safety v1: my first safety-focused suite, to determine how safe a model is today, not if it will take over the world in 2030, but if it actually has real safety issues that could cause problems now
Replying to @claudeai
Opus 5.5 is a major step up from Opus 5, leading on agentic coding, computer use, and knowledge work.
5
1
20
3,026