Fine-tuned specialist vs frontier flagship:
• On contracts, it quoted clauses verbatim 7x as often, at a twentieth of the cost
• On BioRED, a fine-tuned Qwen 9B was right 4x as often
• On NASA ASRS, a 12B model tied on meaning and won on wording
Bigger wasn't better. Think Smaller.
ALT Scoreboard: Overmind specialist vs frontier flagship across legal, biomedical and aviation benchmarks. Right answers: 7x exact quote match (59.6% vs 8.5%), 4x accuracy (54.9% vs 14.1%), +10% wording match (0.295 vs 0.268). Made-up answers: 28x fewer fabricated quotes (0.12% vs 3.36%), 7x fewer invented relationships (6.6% vs 48.2%), 42% fewer invented details (19.4% vs 33.4%). Cost: 20x cheaper per 1,000 questions ($1.03 vs $20.94), 35% cheaper ($17.49 vs $26.73), 14 to 23% fewer words per synopsis.
Across legal, biomedical and aviation benchmarks, our fine-tuned models were 7x more accurate at quoting clauses word for word and 20x cheaper on the legal task.
How we did it:
overmindlab.ai/research/when…