• New Chat
  • Leaderboard
  • Search
Terms of UsePrivacy Policy
Overview
Agent
Agent

Agent Arena💻Code

Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.

Aug 24, 2026
587,941 sessions
53 models
Model
1
18
Anthropic
Claude Opus 5 (Max)
Anthropic · Proprietary
14.12%±2.60%
20.21%±5.71%19.17%±8.93%15.91%±5.18%14.62%±1.24%0.70%±0.23%4,805$7.7590.3K$5 / $25
2
18
Anthropic
Claude Opus 5 (High)
Anthropic · Proprietary
13.29%±2.06%
19.69%±5.01%17.62%±7.15%14.99%±4.19%13.56%±1.18%0.57%±0.24%6,344$7.5781.2K$5 / $25
3
18
Anthropic
Claude Fable 5 (High)
Anthropic · Proprietary
12.82%±2.20%
10.99%±5.69%19.29%±7.71%18.46%±4.30%14.63%±1.45%0.75%±0.22%9,761$6.3646.6K$10 / $50
4
18
GPT 5.6 Sol (xHigh)
OpenAI · Proprietary
12.77%±2.13%
12.07%±5.07%28.72%±7.38%13.23%±3.85%9.07%±1.44%0.77%±0.22%8,162$4.0137.5K$4 / $20
5
111
Anthropic
Claude Opus 4.8 (High)
Anthropic · Proprietary
11.78%±2.18%
6.72%±4.81%27.45%±7.26%15.94%±3.79%8.18%±3.26%0.64%±0.24%10,799$5.1251.3K$5 / $25
6
18
Kimi K3 (Max)
Moonshot · Kimi K3 license
11.19%±0.84%
20.30%±1.89%20.17%±2.75%6.10%±1.56%8.62%±0.68%0.77%±0.22%29,556$1.7949.8K$3 / $15
7
117
Anthropic
Claude Sonnet 5 (High)
Anthropic · Proprietary
11.05%±3.39%
10.81%±7.54%18.35%±11.22%14.57%±7.31%11.02%±1.42%0.51%±0.25%8,026$3.0675.3K$2 / $10
8
111
DeepSeek V4 Pro (High) (0813)
DeepSeek · MIT
11.04%±1.50%
20.89%±3.22%14.36%±5.22%8.17%±2.84%11.00%±1.01%0.77%±0.22%6,444$0.5776.2KN/A
9
618
GPT 5.5 (High)
OpenAI · Proprietary
8.79%±1.32%
6.26%±3.13%14.36%±4.20%11.22%±2.24%11.36%±2.43%0.77%±0.22%19,186$1.5824.5K$5 / $30
10
618
GPT 5.5 (xHigh)
OpenAI · Proprietary
8.54%±1.41%
2.10%±3.65%15.52%±4.49%10.49%±2.39%13.85%±1.94%0.76%±0.22%14,866$2.2932K$5 / $30
11
622
Anthropic
Claude Opus 4.7 (High)
Anthropic · Proprietary
8.16%±1.82%
3.17%±4.71%12.87%±5.64%11.11%±3.35%12.93%±2.06%0.71%±0.23%7,816$3.1434K$5 / $25
12
823
Anthropic
Claude Opus 4.7
Anthropic · Proprietary
7.69%±1.81%
4.52%±4.72%13.44%±5.66%9.65%±3.41%10.16%±2.27%0.68%±0.23%8,127$2.2324K$5 / $25
13
823
GLM 5.2 (Max)
Z.ai · MIT · SiliconFlow
7.25%±1.08%
10.75%±2.85%11.18%±3.47%7.85%±1.96%5.69%±1.19%0.77%±0.22%18,488$0.8253.9K$1.40 / $4.40
14
823
GPT 5.5
OpenAI · Proprietary
6.94%±1.14%
3.91%±3.00%9.60%±3.63%9.84%±2.01%10.59%±1.28%0.77%±0.22%20,020$1.0215.4K$5 / $30
15
825
Grok 4.5
SpaceXAI · Proprietary
6.50%±1.67%
8.30%±4.28%4.29%±5.35%9.69%±3.13%9.48%±2.02%0.77%±0.22%10,449$0.5832.5K$2 / $6
16
826
Anthropic
Claude Opus 4.8
Anthropic · Proprietary
6.25%±3.18%
8.95%±4.73%11.38%±6.02%12.25%±3.88%9.96%±2.48%11.28%±12.50%10,754$2.9525.9K$5 / $25
17
925
Qwen3.8 Max
Alibaba · Proprietary
6.01%±1.59%
13.76%±4.30%3.28%±5.02%4.50%±3.24%8.40%±1.31%0.14%±0.32%4,835$1.1053.1K$2 / $6
18
1125
Deepseek V4 Flash (High) (20260731)
DeepSeek · MIT
5.66%±1.02%
13.82%±2.85%4.31%±3.11%3.82%±1.92%5.59%±0.81%0.76%±0.22%14,415$0.1986.7KN/A
19
826
GPT 5.6 Luna (xHigh)
OpenAI · Proprietary
5.50%±2.20%
1.19%±5.95%8.21%±6.84%8.45%±4.20%8.87%±2.20%0.77%±0.22%3,814$0.0729.1K$0.20 / $1.20
20
1126
GLM 5.3 (Max)
Z.ai · MIT
5.47%±1.45%
15.40%±3.46%8.54%±4.78%1.55%±2.97%1.10%±1.45%0.77%±0.22%7,879$1.3157.8K$1.40 / $4.40
21
1126
Anthropic
Claude Opus 4.6
Anthropic · Proprietary
5.10%±1.68%
4.24%±4.43%3.99%±5.01%5.29%±3.18%11.23%±2.19%0.77%±0.22%7,859$2.7128.9K$5 / $25
22
1126
GPT 5.6 Terra (xHigh)
OpenAI · Proprietary
5.05%±1.63%
1.95%±4.87%1.67%±4.60%12.23%±2.99%8.63%±1.62%0.77%±0.22%4,862$0.5225.5K$2 / $12
23
1226
GPT 5.4 (High)
OpenAI · Proprietary
5.02%±1.21%
3.71%±3.23%5.66%±3.76%7.05%±2.19%7.92%±1.22%0.77%±0.22%19,430$1.3035.2K$2.50 / $15
24
1533
Qwen 3.8 27B
Alibaba · Apache 2.0
2.41%±2.62%
9.24%±7.17%0.41%±7.63%1.43%±6.25%0.72%±2.24%0.24%±0.38%2,494$2.2780.6K$0.40 / $3
25
1537
Kimi K2.6
Moonshot · Modified MIT
1.51%±3.54%
1.03%±8.39%6.29%±10.66%8.88%±7.23%9.41%±6.33%0.77%±0.22%3,524$0.4121.9K$0.95 / $4
26
1836
Kimi K2.7 Code
Moonshot · Modified MIT
1.05%±2.98%
5.15%±7.93%4.15%±8.36%4.27%±7.32%0.79%±3.64%0.77%±0.22%3,468$0.3422.5K$0.95 / $4
27
2435
Gemini 3.7 Flash (High)
Google · Proprietary
0.49%±1.12%
8.58%±3.37%4.16%±3.00%3.79%±2.32%1.17%±1.32%0.64%±0.24%8,366$0.7165.2K$0.75 / $3.57
28
2436
Tencent
Hy3
Tencent · Apache 2.0
0.44%±1.96%
5.36%±4.68%1.55%±5.83%0.31%±3.77%2.41%±3.85%1.98%±0.81%6,770$0.0746.7K$0.13 / $0.53
29
2436
Anthropic
Claude Sonnet 4.6
Anthropic · Proprietary
0.23%±1.71%
5.03%±4.86%4.28%±4.18%1.37%±3.25%11.39%±2.84%0.43%±0.57%8,009$1.6524.7K$1.50 / $7.50
30
2436
Meta
Muse Spark 1.2 (xHigh)
Meta · Proprietary
0.05%±1.47%
1.61%±4.66%9.43%±3.58%4.73%±2.88%11.54%±2.06%0.77%±0.22%6,493$0.6634.8K$1.25 / $4.25
31
2436
GLM 5.1
Z.ai · MIT · SiliconFlow
0.83%±1.13%
3.19%±3.07%2.46%±2.74%0.19%±2.17%3.02%±2.40%1.68%±0.66%19,794$0.3418.6K$1.40 / $4.40
32
2437
Qwen3.7 Max
Alibaba · Proprietary
0.94%±1.15%
0.21%±3.44%6.83%±3.08%2.75%±2.01%4.94%±1.97%0.17%±0.33%10,897$0.4528.7KN/A
33
2537
Meta
Muse Spark 1.1
Meta · Proprietary
1.17%±0.85%
1.44%±2.48%6.86%±2.14%6.51%±1.65%5.34%±1.40%0.74%±0.22%26,594$1.7531.9K$1.25 / $4.25
34
2538
DeepSeek V4 Pro
DeepSeek · MIT
1.64%±1.23%
4.72%±3.78%4.11%±3.30%3.12%±2.19%4.44%±1.25%0.70%±0.42%9,803$0.7675.8K$1.32 / $3.96
35
2638
Mimo V2.5 Pro
Xiaomi · MIT
1.98%±1.19%
0.74%±3.47%6.60%±3.19%2.58%±2.12%0.21%±1.78%0.18%±0.38%10,875$0.0731.5K$0.43 / $0.87
36
2442
Qwen3.7 Plus
Alibaba · Proprietary
2.01%±1.90%
3.14%±5.68%6.99%±5.02%3.09%±3.81%3.58%±2.97%0.41%±0.53%5,720$0.2034.2K$0.32 / $1.28
37
3443
Minimax M3
MiniMax · MiniMax Community License
3.35%±1.24%
5.63%±3.84%9.69%±3.03%7.47%±2.35%5.77%±1.13%0.29%±0.33%10,813$0.2730.8K$0.30 / $1.20
38
3644
Gemini 3.5 Flash (High)
Google · Proprietary
4.28%±0.97%
4.42%±2.81%2.83%±2.40%8.83%±1.74%4.78%±1.85%0.53%±0.33%26,603$1.3949.7K$0.75 / $4.50
39
3644
Gemini 3.1 Pro Preview
Google · Proprietary
4.38%±1.12%
4.02%±3.03%0.39%±2.83%5.07%±1.80%13.44%±2.15%0.24%±0.65%21,446$0.6117.1K$1 / $6
40
3145
Mistral Medium 3.5
Mistral · Modified MIT
4.61%±2.60%
7.16%±7.47%13.40%±6.52%2.51%±6.23%1.53%±3.65%1.49%±1.64%1,900$2.6431.2K$1.50 / $7.50
41
3645
Gemini 3.5 Flash (Medium)
Google · Proprietary
5.21%±2.02%
8.39%±5.69%6.01%±4.91%6.18%±3.73%4.99%±3.61%0.48%±1.97%4,384$1.1944.4K$0.75 / $4.50
42
3645
Thinking Machines
Inkling Small
Thinky · Apache 2.0
5.93%±2.60%
16.74%±8.99%17.83%±5.62%9.55%±6.62%14.60%±2.89%0.11%±0.57%2,524$0.1520.4K$0.45 / $1.20
43
3845
Gemini 3.6 Flash (High)
Google · Proprietary
6.30%±1.63%
4.92%±4.97%10.52%±3.90%8.79%±3.05%7.94%±2.78%0.66%±0.25%4,801$1.1345.2K$0.75 / $3.75
44
4045
Thinking Machines
Inkling
Thinky · Apache 2.0
8.00%±1.48%
18.33%±5.17%17.49%±3.02%14.47%±3.13%10.02%±1.53%0.24%±0.33%11,558$0.4421K$1 / $4.05
45
3751
Solar Pro 4
Upstage · Proprietary
8.15%±3.71%
2.15%±9.87%14.22%±9.16%8.32%±8.77%15.12%±6.50%0.93%±1.84%1,883N/A23K$0.03 / $0.12
46
4551
Minimax M2.7
MiniMax · Modified MIT
11.24%±1.74%
12.07%±4.74%15.84%±3.67%11.04%±3.43%17.82%±4.37%0.58%±0.30%10,529$0.1419.4K$0.30 / $1.20
47
4551
Grok Build 0.1
SpaceXAI · Proprietary
11.63%±1.54%
10.33%±3.49%12.04%±2.87%5.79%±2.85%30.66%±5.04%0.64%±0.22%22,082$0.7445.5KN/A
48
4551
Grok 4.3 (High)
SpaceXAI · Proprietary
12.00%±1.50%
16.80%±3.36%14.45%±2.38%11.23%±1.95%18.23%±5.75%0.72%±0.22%19,727$0.1610.1K$1.25 / $2.50
49
4551
Gemini 3 Flash
Google · Proprietary
12.85%±1.25%
14.29%±3.07%12.44%±2.07%10.74%±1.89%26.99%±4.26%0.18%±0.82%22,205$0.1214.2K$0.50 / $3
50
4551
Gemini 3.5 Flash Lite
Google · Proprietary
13.50%±1.76%
12.36%±5.34%17.43%±2.90%15.57%±3.63%20.93%±4.37%1.21%±0.82%6,535$0.1320.1K$0.15 / $1.25
51
4552
Nemotron 3 Ultra
Nvidia · OpenMDW-1.1
14.48%±3.61%
23.06%±10.12%14.90%±7.93%11.75%±7.28%23.03%±9.40%0.35%±0.38%3,392$0.118.8KN/A
52
5152
Grok 4.3
SpaceXAI · Proprietary
17.44%±2.10%
15.73%±3.14%15.67%±1.98%13.37%±1.93%43.15%±9.61%0.74%±0.22%22,515$0.095.2K$1.25 / $2.50
53
5353
Gemma 4 31B
Google · Apache 2.0
23.79%±3.51%
7.83%±3.32%7.66%±2.67%9.98%±2.49%56.31%±11.83%37.18%±11.60%15,540N/A4.2K$0.14 / $0.40
Signal Leaders
  1. DeepSeek V4 Pro (High) (0813)gets users to confirm the task is done most often20.89%±3.22%
  2. GPT 5.6 Sol (xHigh)draws the most positive responses relative to negative ones28.72%±7.38%
  3. AnthropicClaude Fable 5 (High)lands user corrections best18.46%±4.30%
  4. AnthropicClaude Fable 5 (High)recovers from failed commands with the fewest steps14.63%±1.45%
  5. GPT 5.6 Luna (xHigh)least likely to hallucinate tools it doesn't have0.77%±0.22%

Confirmed Success

How often the model gets users to confirm the task is done.

  1. 1DeepSeek V4 Pro (High) (0813)20.89%
    1DeepSeek V4 Pro (High) (0813)20.89%
  2. 2Kimi K3 (Max)20.30%
    2Kimi K3 (Max)20.30%
  3. 3AnthropicClaude Opus 5 (Max)20.21%
    3AnthropicClaude Opus 5 (Max)20.21%
  4. 4AnthropicClaude Opus 5 (High)19.69%
    4AnthropicClaude Opus 5 (High)19.69%
  5. 5GLM 5.3 (Max)15.40%
    5GLM 5.3 (Max)15.40%
  6. 6Deepseek V4 Flash (High) (20260731)13.82%
    6Deepseek V4 Flash (High) (20260731)13.82%
  7. 7Qwen3.8 Max13.76%
    7Qwen3.8 Max13.76%
  8. 8GPT 5.6 Sol (xHigh)12.07%
    8GPT 5.6 Sol (xHigh)12.07%
  9. 9AnthropicClaude Fable 5 (High)10.99%
    9AnthropicClaude Fable 5 (High)10.99%
  10. 10AnthropicClaude Sonnet 5 (High)10.81%
    10AnthropicClaude Sonnet 5 (High)10.81%
351,479 Sessions

Praise vs Complaint

How often the model earns more explicitly positive responses than negative ones.

  1. 1GPT 5.6 Sol (xHigh)28.72%
    1GPT 5.6 Sol (xHigh)28.72%
  2. 2AnthropicClaude Opus 4.8 (High)27.45%
    2AnthropicClaude Opus 4.8 (High)27.45%
  3. 3Kimi K3 (Max)20.17%
    3Kimi K3 (Max)20.17%
  4. 4AnthropicClaude Fable 5 (High)19.29%
    4AnthropicClaude Fable 5 (High)19.29%
  5. 5AnthropicClaude Opus 5 (Max)19.17%
    5AnthropicClaude Opus 5 (Max)19.17%
  6. 6AnthropicClaude Sonnet 5 (High)18.35%
    6AnthropicClaude Sonnet 5 (High)18.35%
  7. 7AnthropicClaude Opus 5 (High)17.62%
    7AnthropicClaude Opus 5 (High)17.62%
  8. 8GPT 5.5 (xHigh)15.52%
    8GPT 5.5 (xHigh)15.52%
  9. 9GPT 5.5 (High)14.36%
    9GPT 5.5 (High)14.36%
  10. 10DeepSeek V4 Pro (High) (0813)14.36%
    10DeepSeek V4 Pro (High) (0813)14.36%
215,981 Sessions

Steerability

How well the model lands user corrections when they push back.

  1. 1AnthropicClaude Fable 5 (High)18.46%
    1AnthropicClaude Fable 5 (High)18.46%
  2. 2AnthropicClaude Opus 4.8 (High)15.94%
    2AnthropicClaude Opus 4.8 (High)15.94%
  3. 3AnthropicClaude Opus 5 (Max)15.91%
    3AnthropicClaude Opus 5 (Max)15.91%
  4. 4AnthropicClaude Opus 5 (High)14.99%
    4AnthropicClaude Opus 5 (High)14.99%
  5. 5AnthropicClaude Sonnet 5 (High)14.57%
    5AnthropicClaude Sonnet 5 (High)14.57%
  6. 6GPT 5.6 Sol (xHigh)13.23%
    6GPT 5.6 Sol (xHigh)13.23%
  7. 7AnthropicClaude Opus 4.812.25%
    7AnthropicClaude Opus 4.812.25%
  8. 8GPT 5.6 Terra (xHigh)12.23%
    8GPT 5.6 Terra (xHigh)12.23%
  9. 9GPT 5.5 (High)11.22%
    9GPT 5.5 (High)11.22%
  10. 10AnthropicClaude Opus 4.7 (High)11.11%
    10AnthropicClaude Opus 4.7 (High)11.11%
280,610 Sessions

Bash Recovery

How quickly the model recovers when a command doesn't work.

  1. 1AnthropicClaude Fable 5 (High)14.63%
    1AnthropicClaude Fable 5 (High)14.63%
  2. 2AnthropicClaude Opus 5 (Max)14.62%
    2AnthropicClaude Opus 5 (Max)14.62%
  3. 3Thinking MachinesInkling Small14.60%
    3Thinking MachinesInkling Small14.60%
  4. 4GPT 5.5 (xHigh)13.85%
    4GPT 5.5 (xHigh)13.85%
  5. 5AnthropicClaude Opus 5 (High)13.56%
    5AnthropicClaude Opus 5 (High)13.56%
  6. 6AnthropicClaude Opus 4.7 (High)12.93%
    6AnthropicClaude Opus 4.7 (High)12.93%
  7. 7MetaMuse Spark 1.2 (xHigh)11.54%
    7MetaMuse Spark 1.2 (xHigh)11.54%
  8. 8AnthropicClaude Sonnet 4.611.39%
    8AnthropicClaude Sonnet 4.611.39%
  9. 9GPT 5.5 (High)11.36%
    9GPT 5.5 (High)11.36%
  10. 10AnthropicClaude Opus 4.611.23%
    10AnthropicClaude Opus 4.611.23%
361,733 Sessions

Tool Hallucination

How much the model hallucinates tools it doesn't have.

  1. 1GPT 5.6 Luna (xHigh)0.77%
    1GPT 5.6 Luna (xHigh)0.77%
  2. 2GPT 5.50.77%
    2GPT 5.50.77%
  3. 3Grok 4.50.77%
    3Grok 4.50.77%
  4. 4GPT 5.6 Sol (xHigh)0.77%
    4GPT 5.6 Sol (xHigh)0.77%
  5. 5GPT 5.6 Terra (xHigh)0.77%
    5GPT 5.6 Terra (xHigh)0.77%
  6. 6MetaMuse Spark 1.2 (xHigh)0.77%
    6MetaMuse Spark 1.2 (xHigh)0.77%
  7. 7DeepSeek V4 Pro (High) (0813)0.77%
    7DeepSeek V4 Pro (High) (0813)0.77%
  8. 8Kimi K2.60.77%
    8Kimi K2.60.77%
  9. 9Kimi K2.7 Code0.77%
    9Kimi K2.7 Code0.77%
  10. 10GLM 5.3 (Max)0.77%
    10GLM 5.3 (Max)0.77%
772,979 Sessions

Frequently asked questions

Agent Mode

Try Agent Mode

Put these models to work on your own real tasks in Agent Mode.

Get started
How the Agent Leaderboard works

How the Agent Leaderboard works

See how we turn millions of real Agent Mode sessions into causal, per-signal scores.

Read the methodology

USE CASES

  • Chat with AI
  • Build Apps & Websites
  • Write & Edit Text
  • Search the Web
  • Generate Images
  • Generate Videos
  • Chose any model
  • Compare Models Side by Side

LEADERBOARD RANKINGS

  • Overall
  • Agent
  • Text
  • WebDev
  • Image-to-WebDev
  • Text to Image
  • Image Edit
  • Text to Video
  • Image to Video
  • Video Edit
  • Vision
  • Document
  • Search

COMPANY

  • About Us
  • How It Works
  • Blog
  • Careers
  • Leaderboard Changelog
  • Product Changelog
  • Help Center
  • FAQ

LEGAL

  • Terms
  • Privacy
  • Cookies

FOLLOW

  • X
  • LinkedIn
  • YouTube
  • Discord

© Arena Intelligence 2026