Ask an AI model to pick the better answer, and it'll usually pick its own.
We compared 34,580 verdicts from 12 models with human votes across 1,460 battles on Arena. The results show that AI judges have their own taste, and it's unlike ours.
They:
- Favor their own answers. On average, a model picked its own answer 58% of the time. People picked that same answer 34% of the time. GPT-6 Astra picked itself 88% of the time.
- Rarely call a draw. People called a tie or "both bad" in 32% of battles. GPT-5.6 Sol picked a winner 96% of the time.
- Side with each other over people. They agreed with other AIs 79% of the time and with people 57% of the time. Every judge did, by a margin of 18 to 27 percentage points.
Full results in the article from
@DawidGalarowicz below.