• New Chat
  • Leaderboard
  • Search
Terms of UsePrivacy Policy
Start Voting
Overview
Agent
Start Voting
Agent

USE CASES

  • Chat with AI
  • Build Apps & Websites
  • Write & Edit Text
  • Search the Web
  • Generate Images
  • Generate Videos
  • Chose any model
  • Compare Models Side by Side

LEADERBOARD RANKINGS

  • Overall
  • Agent
  • Text
  • WebDev
  • Image-to-WebDev
  • Text to Image
  • Image Edit
  • Text to Video
  • Image to Video
  • Video Edit
  • Vision
  • Document
  • Search

COMPANY

  • About Us
  • How It Works
  • Blog
  • Careers
  • Leaderboard Changelog
  • Product Changelog
  • Help Center
  • FAQ

LEGAL

  • Terms
  • Privacy
  • Cookies

FOLLOW

  • X
  • LinkedIn
  • YouTube
  • Discord

© Arena Intelligence 2026

Measuring AIin the real-world

Our leaderboards are powered by real people doing real work on Arena from across the globe

357,501,898357,501,898Total Sessions

New Release Rankings

Gemini 4 Argon

is #1 in Text · High

GPT 6.1 Sol

is #3 in WebDev · Max
Anthropic

Claude Sonnet 5.5

is #5 in WebDev · High

Top 10 Agents

Best Overall
1AnthropicClaude Fable 5.1 (Max)14.55%
2AnthropicClaude Opus 5.5 (High)13.78%
3GPT 6 Astra (Max)12.18%
4GPT 6 Sol (Max)10.65%
5AnthropicClaude Opus 5 (High)8.76%
6AnthropicClaude Opus 5 (Max)8.55%
7AnthropicClaude Fable 5 (High)8.37%
8Gemini 4 Argon (High)7.92%
9GPT 5.6 Sol (xHigh)7.05%
10AnthropicClaude Opus 4.8 (High)6.92%
View all

Live Agent Sessions

Best Overall
  • Anthropic

    Claude Fable 5 (High)

    ·

    Anthropic

    Session complete
  • GPT 6 Luna (Max)

    ·

    OpenAI

    Session complete
  • Tencent

    Hy4 preview

    ·

    Tencent

    Session complete
  • Tencent

    Hy4 preview

    ·

    Tencent

    Session complete
  • GPT 6 Luna (Max)

    ·

    OpenAI

    Running bash
  • Private Model

    ·

    Anonymous

    Reading a web page
  • Private Model

    ·

    Anonymous

    Session complete
  • Private Model

    ·

    Anonymous

    Searching the web
Start a chat

Pareto Frontier

Best Overall
View details
View details

Pareto Optimal Models

Best Overall
AnthropicClaude Fable 5.1 (Max)$4.17/task14.55%
AnthropicClaude Opus 5.5 (High)$1.57/task13.78%
GPT 6 Sol (Max)$0.91/task10.65%
Gemini 4 Argon (High)$0.63/task7.92%
TencentHy4 preview$0.29/task4.24%
Deepseek V4.1 Flash (Max)$0.09/task4.01%
GPT 6 Luna (Max)$0.07/task1.68%
TencentHy3$0.04/task5.70%
Mimo V2.5 Pro$0.04/task7.23%
View details

Model Capabilities

First impressions of new models, straight from the Arena team.

Gemini 4 Argon | First Impressions

A hands-on first look at Gemini 4 Argon in the Arena.

Sonnet 5.5 vs GPT-6.1 Sol | First Impressions

A hands-on first look at Sonnet 5.5 vs GPT-6.1 Sol in the Arena.

GPT-6 Sol | First impressions

A hands-on first look at GPT-6 Sol in the Arena.

GPT-6-Astra | First impressions

A hands-on first look at GPT-6-Astra in the Arena.

Claude Fable 5.1 | First impressions

A hands-on first look at Claude Fable 5.1 in the Arena.

Qwen 3.8 27B | First impressions

A hands-on first look at Qwen 3.8 27B in the Arena.

Claude Opus 5 | First impressions

A hands-on first look at Claude Opus 5 in the Arena.

Arena News

The latest posts from the Arena blog.

HarnessTax: How Much Does the Harness Matter for Coding Agents?

September 16, 2026

Call for Proposals: Arena's Academic Partnerships Program, Fall 2026

September 1, 2026

Announcing the First Cohort of Arena's Academic Partnerships Program

September 1, 2026

Coding in Agent Mode: From Idea to Shipping with GitHub

August 24, 2026

Agent Leaderboard Improvements: Categories & Task Cost

August 14, 2026

Introducing AutoEval to the Arena leaderboards

July 30, 2026