• New Chat
  • Leaderboard
  • Search
Terms of UsePrivacy Policy
Start Voting
Overview
Agent
Start Voting
Agent

USE CASES

  • Chat with AI
  • Build Apps & Websites
  • Write & Edit Text
  • Search the Web
  • Generate Images
  • Generate Videos
  • Chose any model
  • Compare Models Side by Side

LEADERBOARD RANKINGS

  • Overall
  • Agent
  • Text
  • WebDev
  • Image-to-WebDev
  • Text to Image
  • Image Edit
  • Text to Video
  • Image to Video
  • Video Edit
  • Vision
  • Document
  • Search

COMPANY

  • About Us
  • How It Works
  • Blog
  • Careers
  • Leaderboard Changelog
  • Product Changelog
  • Help Center
  • FAQ

LEGAL

  • Terms
  • Privacy
  • Cookies

FOLLOW

  • X
  • LinkedIn
  • YouTube
  • Discord

© Arena Intelligence 2026

Measuring AIin the real-world

Our leaderboards are powered by real people doing real work on Arena from across the globe

356,334,944356,334,944Total Sessions

New Release Rankings

Mimo V2.6 Flash

is #45 in Vision
Stepfun

Step 5

is #29 in WebDev · Preview High
Recraft

Recraft V4.1 Flash

is #55 in Text-to-Image

Top 10 Agents

Best Overall
1AnthropicClaude Fable 5.1 (Max)14.06%
2AnthropicClaude Opus 5.5 (High)11.84%
3GPT 6 Astra (Max)10.36%
4AnthropicClaude Opus 5 (Max)9.54%
5AnthropicClaude Opus 5 (High)9.38%
6GPT 6 Sol (Max)8.80%
7AnthropicClaude Fable 5 (High)7.99%
8AnthropicClaude Opus 4.8 (High)6.87%
9GPT 5.6 Sol (xHigh)6.26%
10AnthropicClaude Sonnet 5 (High)4.80%
View all

Live Agent Sessions

Best Overall
  • GPT 6 Luna (Max)

    ·

    OpenAI

    Session complete
  • Gemini 3.7 Flash (High)

    ·

    Google

    Writing a file
  • Deepseek V4.1 Flash (Max)

    ·

    DeepSeek

    Session complete
  • Anthropic

    Claude Opus 4.8 (High)

    ·

    Anthropic

    Running bash
  • GLM 5.3 Flash

    ·

    Z.ai

    Session complete
  • Anthropic

    Claude Opus 5 (High)

    ·

    Anthropic

    Reading files
  • GPT 6 Luna (Max)

    ·

    OpenAI

    Running bash
  • Anthropic

    Claude Opus 5.5 (High)

    ·

    Anthropic

    Running bash
Start a chat

Pareto Frontier

Best Overall
View details
View details

Pareto Optimal Models

Best Overall
AnthropicClaude Fable 5.1 (Max)$3.48/task14.06%
AnthropicClaude Opus 5.5 (High)$1.37/task11.84%
GPT 6 Sol (Max)$0.81/task8.80%
AnthropicClaude Sonnet 5 (High)$0.74/task4.80%
TencentHy4 preview$0.21/task4.10%
Deepseek V4.1 Flash (Max)$0.07/task3.86%
GPT 6 Luna (Max)$0.05/task1.59%
TencentHy3$0.04/task5.75%
Mimo V2.5 Pro$0.03/task6.61%
View details

Model Capabilities

First impressions of new models, straight from the Arena team.

GPT-6 Sol | First impressions

A hands-on first look at GPT-6 Sol in the Arena.

GPT-6-Astra | First impressions

A hands-on first look at GPT-6-Astra in the Arena.

Claude Fable 5.1 | First impressions

A hands-on first look at Claude Fable 5.1 in the Arena.

Qwen 3.8 27B | First impressions

A hands-on first look at Qwen 3.8 27B in the Arena.

Claude Opus 5 | First impressions

A hands-on first look at Claude Opus 5 in the Arena.

Kimi K3 | First impressions

A hands-on first look at Kimi K3 in the Arena.

Arena News

The latest posts from the Arena blog.

HarnessTax: How Much Does the Harness Matter for Coding Agents?

September 16, 2026

Call for Proposals: Arena's Academic Partnerships Program, Fall 2026

September 1, 2026

Announcing the First Cohort of Arena's Academic Partnerships Program

September 1, 2026

Coding in Agent Mode: From Idea to Shipping with GitHub

August 24, 2026

Agent Leaderboard Improvements: Categories & Task Cost

August 14, 2026

Introducing AutoEval to the Arena leaderboards

July 30, 2026