Millions are spent training models for specific capabilities, then judging them with: “Which output do you prefer?” 💸
We need more granularity + transparency. So here comes
benchmark.ai.
First benchmark: SVG generation. 42 models, 1.9M+ human judgements.
Prompts, outputs, match-ups & methodology are public. What benchmark do you want to see next?