

- Published on 26 May 2023
- Last updated on 13 Jun 2025
- Reading Time: 8 minutes
Chatbot Arena Leaderboard Updates (Week 4)
A new Elo rating leaderboard based on the 27K anonymous voting data collected in the wild between April 24 and May 22, 2023 is released in Table 1 below.
In this update, we are excited to welcome the following models joining the Chatbot Arena:
- Google PaLM 2, chat-tuned with the code name chat-bison@001 on Google Cloud Vertex AI
- Anthropic Claude-instant-v1
- MosaicML MPT-7B-chat
- Vicuna-7B
We provide a Google Colab notebook to analyze the voting data, including the computation of the Elo ratings. You can also try the voting demo.

Win Fraction Matrix
The win fraction matrix of all model pairs is shown in Figure 1.

Overview
Google PaLM 2
Google’s PaLM 2 is one of the most significant models announced since our last leaderboard update. We added the PaLM 2 Chat to the Chatbot Arena via the Google Cloud Vertex AI API. The model is chat-tuned under the code name chat-bison@001.
In the past two weeks, PaLM 2 has competed for around 1.8k anonymous battles with the other 16 chatbots, currently ranked 6th on the leaderboard. It ranks above all other open-source chatbots, except for Vicuna-13B, whose Elo is 12 scores higher than PaLM 2 (Vicuna 1054 vs. PaLM 2 1042) which in terms of ELO rating is nearly a virtual tie. We noted the following interesting results from PaLM 2’s Arena data.
PaLM 2 is better when playing against the top 4 players, i.e., GPT-4, Claude-v1, ChatGPT, Claude-instant-v1, and it also wins 53% of the plays with Vicuna, but worse when playing against weaker players. This can be seen in Figure 1 which shows the win fraction matrix. Among all battles PaLM 2 has participated in, 21.6% were lost to a chatbot that is not one of GPT-4, Claude-v1, GPT-3.5-turbo, Claude-instant-v1. For reference, another proprietary model GPT-3.5-turbo only loses 12.8% of battles to those chatbots.In short, we find that the current PaLM 2 version available at Google Cloud Vertex API has the following deficiencies when compared to other models we have evaluated:












