

- Published on 8 Dec 2023
- Last updated on 13 Jun 2025
- Reading Time: 10 minutes
Chatbot Arena - New models & Elo system update
Welcome to our latest update on the Chatbot Arena, our open evaluation platform to test the most advanced LLMs. We’re excited to share that over 130,000 votes that are now collected to rank the most capable 40+ models! In this blog post, we’ll cover the results of several new models:
- Tulu-2-DPO-70B and Yi-34B-Chat are the new SoTA open models
- Mistral-based 7B models (OpenChat, OpenHermes-2.5, Starling-7B) show promising performance
We also present our findings from differentiating versions of proprietary models (e.g., GPT-4 => GPT-4-0314, GPT-4-0613), and the transition from the online Elo system to the Bradley-Terry model, which gives us significantly more stable ratings and precise confidence intervals.
Let’s dive into it!
Introducing new models
LLM has become smarter than ever and it’s been a real challenge to evaluate them properly. Traditional benchmarks such as MMLU have been useful, but they may fall short in capturing the nuance of human preference and open-ended nature of real-world conversations. We believe deploying chat models in the real-world to get feedback from users produces the most direct signals. This led to the Chatbot Arena launch in May. Since then, the open-source community has taken off. Over the past few months, we have deployed more than 45 models in Arena and we’ve collected over 130,000 valid votes from our users. We believe such a scale covers a diverse range of use cases which bring us useful insights to understand how these models work in real-world scenarios.
In November, we added record-breaking nine new models with sizes ranging from 7B to 70B, as well as proprietary ones, and gathered over new 25,000 votes for them. Excitingly, we are now seeing the gap between proprietary and open models narrowing. New models such as Tulu-2-DPO-70B and Yi-34B-Chat have been leading the open space, delivering close to gpt-3.5 performance.

On the other hand, 7B models have also shown significant improvements. Fine-tuning the 7B Mistral model has led to Zephyr, OpenChat-3.5, Starling-lm-7b-alpha, and OpenHermes-2.5-Mistral-7b which all demonstrate impressive performance despite smaller scale. Shoutout to the open-source community pushing limits! On the other hand, to understand how freshness and grounded information help LLMs in answering user queries, we also bring Perplexity AI’s online LLMs to Arena. We have collected over 1500 votes for PPLX-70B-Online and the preliminary results show great potential. Congrats to all the teams and we look forward to seeing more models in the future!Please find the latest leaderboard or try to chat with 20+ models! We also prepare a to reproduce all the calculation of Elo ratings and confidence intervals.
















