I rely on AI heavily to maintain my OSINT analysis enterprise.
For defense readers, you can skip to the first reply to see warfare and Ukraine implications.
For those interested in how Composer 2.5 holds its own against the Frontier, keep reading.
We don't always need Frontier models when many older models get the job done.
Transforming equations and data collection to code that runs on VMs 24/7 takes time and effort
Just as I benchmark my oil and interception models, I also rely on benchmarks for my AI tools. As I do this on my own budget, costs matter. Many benchmarks use hypothetical problems that don't relate to normal work.
@morganlinton has made a benchmark
@VulcanBench that does a great job of testing models against real world stressing problems.
Recently he compared Fable 5.1, Astra, and Opus 5.5. I decided to compare it to Composer 2.5, an older but very well developed and economical mode by
@cursor_ai . It held is own and was able to accomplish many of the tasks and be cost competitive.
This is why. I want to start by saying Morgan made some amazing tests and his prompts are incredibly well written. Looking at the results for each of the 23 models shows that for shorter tasks, Composer does a great job and is either the cheapest or 2nd cheapest option for about half. For the other half it is more expensive. This is simply a case where it cannot support tasks of that length. Too long and too many steps for the context window.
This is where frontier models shine and why many benchmarks undersell their capabilities versus tailored coding models like many of the Chinese ones.
However, many of these frontier models are great at breaking tasks into smaller chunks so strong, cheap, but older models like Composer 2.5 can handle.
I have used Grok 4.6/4.7, Fable 5/5.1, Sol 5.6, and Opus 5.5 to break down tasks. When paired with Composer, it is like getting those frontier models at a very steep discount.
For cost conscious researcher like myself, most of what we need are Lunas and Sonnets and Composers led by frontier models.
@grok please keep your brother-in-law, Composer 2.5, in the family.