Do blind battles reveal any bias?
Astra picked itself in 88% of its battles, Fable in 72%.
When Arena users enter Battle Mode, their prompt goes to two randomly selected models. They then pick the better answer, call a tie, or say both are bad. We gave the same task to AI judges, without telling them which models wrote the answers.

In battles where one answer was its own, Astra called its own answer better 88% of the time and picked the opponent's only 2% of the time. The people who submitted the prompts picked Astra's answer 29% of the time, the opponent's 36%, and called a tie or "both bad" in the rest. Fable rated its own answer better in 72% of battles, compared with 38% for people.
This self-preference becomes a clear pattern as we consider additional models. Across 12 LLMs, the models backed themselves about 70% more often than humans did. On average, a model picked its own answer 58% of the time, while people picked that same answer 34% of the time.
If we set ties aside, the top self-fans look even more one-sided.

Could self-preference extend to model siblings?
OpenAI judges favor OpenAI models by 37 points versus 6 for everyone else.

To see whether this preference extends to models from the same family, we looked at how every judge rated every author, not just itself.
Our results show that this family loyalty exists. Astra, Sol, and Luna rated other OpenAI models 37 points above people. The other nine judges rated them just 6 points above, and the OpenAI trio rated everyone else 15 points below people. Fable and Opus showed a milder version, with +28 for each other against +15 from other judges. The DeepSeek models were the exception, rating each other 16 points below people, close to the other judges' -20.
There was no national loyalty, though. Chinese judges were 6 points more generous than people to American answers, and 6 points harsher on answers from other Chinese labs. American judges matched people on other American labs' answers, and were 12 points harsher on Chinese ones.
Beyond family, the judges mostly agreed with each other. For example, all ten outside judges liked Fable more than people did, by 20 points on average. Similarly, every outside judge was more generous than people to Kimi. That points to a taste the models share, not just family loyalty.
Do AI judges hesitate as much as humans?
Sol picked a winner 96% of the time, compared with 68% for people.

We also found that models rarely call it a draw. People called a tie or "both bad" in 32% of battles. Every model was more decisive: Sol picked a winner 96% of the time, Luna 94%, and Gemini 93%. DeepSeek V4 Flash was the most human-like, at 71%.
Do AI judges have their own taste?
AI judges agree with each other, not with us.

Finally, to get insight into model taste, we asked two questions of each judge's verdict in every battle. First, did it pick the same answer as the person who voted? Second, did it pick the same answer as the other AI judges? If the models were simply tracking what people like, the two numbers would be similar.
They weren't. The average judge sided with the other models 79.4% of the time, but with the person only 56.9% of the time. GLM had the biggest gap: 85.2% with the other AIs, 58.5% with the person. Every judge sided with the AIs more often, by 18 to 27 points.
Even their unanimous verdicts were shaky. In the 393 battles where every AI judge picked the same answer, the person agreed only 60.6% of the time.
Astra, Sol, and Luna were the only judges that agreed with people less than half the time (48.2%, 48.7%, and 49.6%). But they also agreed least with the other AIs - they weren't following the crowd either. Their OpenAI-flavored taste is their own.
Before you trust an AI judge
If you use models to judge models:
- Don't treat the author as a neutral reviewer. It will likely favor itself.
- Don't assume judges from other labs speak for your users. They can agree with each other and still differ from people.
- Expect judges to force a winner where people would call it close, and track how often each judge declines to choose.
- Ideally, check verdicts against real human preferences or expert views.
Methodology
Data
We used 1,460 unique English-language Text Arena battles from June 1 - September 23, 2026.
Approach
All 12 models were asked to judge every battle. Each judge saw every battle twice, with the answer order swapped.
Each battle counted once per judge:
- Same answer both times: that answer is the judge's winner.
- A winner once and a tie or "both bad" the other time: the winner counts.
- A different answer each time: half a pick in the self-preference and family charts, and no pick in the shared-taste chart.
- No winner either time: the battle is left out of the self-preference and family charts.
Of the 35,040 planned verdicts, 34,580 were valid. The rest were unusable replies, content-filter blocks, or failed calls.










