

- Published on 23 Jun 2023
- Last updated on 13 Jun 2025
- Reading Time: 8 minutes
Chatbot Arena Leaderboard Updates (Week 8)
Introducing MT-Bench and Vicuna-33B
In this blog post, we share the latest update on Chatbot Arena leaderboard, which now includes more open models and three metrics:
- Chatbot Arena Elo, based on 42K anonymous votes from Chatbot Arena using the Elo rating system.
- MT-Bench score, based on a challenging multi-turn benchmark and GPT-4 grading, proposed and validated in our Judging LLM-as-a-judge paper.
- MMLU, a widely adopted benchmark.
Furthermore, we’re excited to introduce our new series of Vicuna-v1.3 models, ranging from 7B to 33B parameters, trained on an extended set of user-shared conversations. Their weights are now available.
Updated Leaderboard and New Models

Welcome to try the Chatbot Arena voting demo. Keep in mind that each benchmark has its limitations. Please consider the results as guiding references. See our discussion below for more technical details.
Evaluating Chatbots with MT-bench and Arena
Motivation
While several benchmarks exist for evaluating Large Language Model’s (LLM) performance, such as , , and , we noticed that these benchmarks might fall short when assessing LLMs’ human preferences. Traditional benchmarks often test LLMs on close-ended questions with concise outputs (e.g., multiple choices), which do not reflect the typical use cases of LLM-based chat assistants.To fill this gap, in this leaderboard update, in addition to the Chatbot Arena Elo system, we add a new benchmark: MT-Bench.












