Skip to main content.
AllNewsProductResearch
Text Arena and Search Arena scores with the factuality adjustment on

Factuality in the Arena

Factuality remains one of the most persistent questions users face when using AI models. Today, we are launching a leaderboard that ranks models not only by human preference, but also by the factual accuracy of their responses.

Read Article
Research
Product
Arena Team—14 Jul 2026
Agent Arena: Causal Evaluation of Agents in the Real World
Read Article

Agent Arena: Causal Evaluation of Agents in the Real World

Agents are increasingly doing real work. The resulting task distribution has greatly expanded. We desire an agent evaluation that scales along with usage and capability.

Product
Research
Arena Team—4 Jun 2026
Arena Leaderboard Dataset
Read Article

Arena Leaderboard Dataset

For almost three years, Arena has been publishing leaderboards covering frontier AI capabilities across 10 arenas, dozens of categories, and hundreds of models, and today we're releasing the entire history of those leaderboards as a public-access dataset.

Research
Arena Team—2 Apr 2026
Supporting Independent Research in AI Evaluation
Read Article

Supporting Independent Research in AI Evaluation

Arena’s Academic Partnerships Program provides funding and support for independent research advancing the scientific foundations of AI evaluation.

Research
News
Arena Team—10 Feb 2026
Introducing Max
Read Article

Introducing Max

Today we are releasing Max, Arena's model router powered by our community’s 5+ million real-world votes. Max acts as an intelligent orchestrator—it routes each user prompt to the most capable model for that specific prompt.

Research
Arena Team—4 Feb 2026
Arena-Rank: Open Sourcing the Leaderboard Methodology
Read Article

Arena-Rank: Open Sourcing the Leaderboard Methodology

Building community trust with open science is critical for the development of AI and its alignment with the needs and preferences of all users. With that in focus, we’re delighted to publish Arena-Rank, an open-source Python package for ranking that powers the LMArena leaderboard!

Research
Arena Team—18 Dec 2025
Studying the Frontier: Arena Expert
Read Article

Studying the Frontier: Arena Expert

Arena Expert is a great way to differentiate between frontier models. In this analysis, we compare how models perform on 'general' vs 'expert' prompts, focusing on 'thinking' vs 'non-thinking' models.

Research
Arena Team—4 Dec 2025

Subscribe
to Arena news and research

Insights at the frontier of AI.

Invalid email address
Arena's Ranking Method
Read Article

Arena's Ranking Method

Since launching the platform, developing a rigorous and scientifically grounded evaluation methodology has been central to our mission. A key component of this effort is providing proper statistical uncertainty quantification for model scores and rankings. To that end, we have always reported confidence intervals alongside Arena scores and surfaced any ties in the rankings that those intervals imply.

Research
News
Arena Team—14 Nov 2025
Arena Expert and Occupational Categories
Read Article

Arena Expert and Occupational Categories

The next frontier of large language model (LLM) evaluation lies in understanding how models perform when challenged by expert-level problems, drawn from real work, across diverse disciplines.

Research
Arena Team—5 Nov 2025
Re-introducing Vision Arena Categories
Read Article

Re-introducing Vision Arena Categories

Since we first introduced categories over two years ago, and Vision Arena last year, the AI evaluation landscape has evolved. New categories have been added, existing ones have been updated, and the leaderboards they power are becoming more insightful with each round of community input.

Research
Arena Team—3 Oct 2025
Introducing BiomedArena.AI: Evaluating LLMs for Biomedical Discovery
Read Article

Introducing BiomedArena.AI: Evaluating LLMs for Biomedical Discovery

LMArena is honored to partner with the team at DataTecnica to advance the expansion of BiomedArena.ai: a new domain-specific evaluation track.

Research
Arena Team—19 Aug 2025
A Deep Dive into Recent Arena Data
Read Article

A Deep Dive into Recent Arena Data

Today, we're excited to release a new dataset of recent battles from LMArena! The dataset contains 140k conversations from the text arena.

Research
Arena Team—31 Jul 2025
Does Sentiment Matter Too?
Read Article

Does Sentiment Matter Too?

Introducing Sentiment Control: Disentangling Sentiment and Substance

Research
Arena Team—22 Apr 2025
Try Arena
Try Arena
Leaderboard Rankings
Overall
Agent
Text
WebDev
Image-to-WebDev
Text to Image
Image Edit
Text to Video
Image to Video
Video Edit
Vision
Document
Search
Use Cases
Chat with AI
Build Apps & Websites
Write & Edit Text
Search the Web
Generate Images
Generate Videos
Chose any Model
Compare Models Side by Side
Complete Multi-step Tasks
Company
About Us
How It Works
Careers
Changelog
Help Center
FAQ
Blog
Follow
X
LinkedIn
YouTube
Discord
TermsPrivacyCookies
Try Arena
Try Arena

Ⓒ 2026 Arena Intelligence Inc.