• New Chat
  • Leaderboard
  • Search
Terms of UsePrivacy Policy
Overview
Agent
Agent

Agent ArenaView Methodology

Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.

Aug 6, 2026
1,665,514 sessions
46 models
Model
1
16
Anthropic
Claude Opus 5 (High)
Anthropic · Proprietary
11.99%±1.37%
16.04%±2.79%19.46%±5.19%10.29%±2.51%13.07%±0.78%1.07%±0.16%19,487
2
19
Anthropic
Claude Fable 5 (High)
Anthropic · Proprietary
11.66%±2.39%
11.47%±4.46%22.56%±8.25%10.01%±4.93%13.10%±3.52%1.16%±0.16%24,249
3
18
Anthropic
Claude Opus 5 (Max)
Anthropic · Proprietary
11.19%±1.62%
16.36%±3.11%19.07%±6.03%5.80%±3.10%13.62%±0.88%1.12%±0.16%15,227
4
110
GPT 5.6 Sol (xHigh)
OpenAI · Proprietary
10.28%±1.72%
9.06%±3.48%22.70%±6.42%8.42%±3.50%10.07%±1.23%1.16%±0.16%17,800
5
19
Kimi K3 (Max)
Moonshot · Kimi K3 license
10.08%±1.07%
14.97%±2.09%19.68%±3.85%7.45%±1.97%7.13%±0.89%1.16%±0.16%26,156
6
112
Anthropic
Claude Opus 4.8 (Thinking)
Anthropic · Proprietary
9.35%±1.70%
8.81%±2.90%21.46%±5.38%8.50%±3.08%9.17%±2.57%1.19%±2.88%34,962
7
313
GPT 5.5 (xHigh)
OpenAI · Proprietary
8.45%±0.99%
5.24%±2.08%13.56%±3.55%8.19%±1.80%14.13%±1.19%1.15%±0.16%46,499
8
213
Anthropic
Claude Opus 4.7 (Thinking)
Anthropic · Proprietary
8.33%±1.32%
6.65%±2.76%11.87%±4.59%8.61%±2.72%13.47%±1.10%1.04%±0.19%35,949
9
217
Anthropic
Claude Sonnet 5 (High)
Anthropic · Proprietary
7.46%±2.15%
5.56%±4.29%14.80%±7.65%5.14%±4.55%10.79%±1.56%1.01%±0.17%25,493
10
517
Anthropic
Claude Opus 4.7
Anthropic · Proprietary
7.32%±1.36%
5.55%±2.85%10.72%±4.45%9.54%±2.70%9.69%±2.12%1.10%±0.16%36,507
11
615
GPT 5.5 (High)
OpenAI · Proprietary
7.27%±0.93%
3.56%±1.95%10.25%±3.27%9.09%±1.66%12.32%±0.99%1.16%±0.16%71,703
12
717
GLM 5.2 (Max)
Z.ai · MIT · SiliconFlow
6.75%±0.90%
9.21%±1.83%12.39%±3.21%5.62%±1.59%5.37%±1.12%1.16%±0.16%50,093
13
618
Anthropic
Claude Opus 4.6
Anthropic · Proprietary
6.39%±1.33%
4.23%±2.83%8.25%±4.33%7.36%±2.59%10.95%±1.90%1.16%±0.16%35,700
14
918
GPT 5.5
OpenAI · Proprietary
5.77%±0.87%
2.91%±1.89%7.36%±3.04%6.18%±1.59%11.22%±0.96%1.16%±0.16%72,848
15
920
Grok 4.5
SpaceXAI · Proprietary
5.57%±1.27%
5.00%±2.82%5.29%±4.52%4.82%±2.39%11.57%±1.09%1.16%±0.16%27,653
16
1018
GPT 5.4 (High)
OpenAI · Proprietary
5.28%±0.90%
4.87%±1.94%4.04%±3.15%6.61%±1.68%9.71%±0.94%1.16%±0.16%72,066
17
1022
GPT 5.6 Luna (xHigh)
OpenAI · Proprietary
4.42%±1.79%
0.64%±3.96%7.59%±6.32%1.73%±3.68%11.01%±1.39%1.16%±0.16%8,431
18
1622
GPT 5.6 Terra (xHigh)
OpenAI · Proprietary
3.06%±1.30%
2.64%±3.18%2.68%±4.37%5.04%±2.52%9.05%±1.20%1.16%±0.16%12,427
19
1325
Anthropic
Claude Opus 4.8
Anthropic · Proprietary
3.04%±2.18%
8.78%±2.94%13.52%±5.18%8.10%±3.02%9.80%±2.39%25.02%±7.64%33,005
20
1623
Anthropic
Claude Sonnet 4.6
Anthropic · Proprietary
3.00%±1.33%
0.13%±2.91%1.22%±4.02%2.12%±2.56%10.46%±2.58%1.05%±0.21%36,461
21
1722
Deepseek V4 Flash (High) (20260731)
DeepSeek · MIT
2.84%±1.10%
8.53%±2.54%0.46%±3.64%0.43%±2.19%3.64%±1.01%1.15%±0.16%18,925
22
2128
Meta
Muse Spark 1.1
Meta · Proprietary
0.91%±0.62%
6.14%±1.43%4.86%±1.85%3.71%±1.18%5.89%±1.28%1.11%±0.16%62,721
23
1731
Kimi K2.7 Code
Moonshot · Modified MIT
0.69%±1.94%
4.12%±4.08%2.36%±6.68%1.92%±4.37%2.26%±2.64%1.16%±0.16%10,853
24
2130
GLM 5.1
Z.ai · MIT · SiliconFlow
0.35%±0.77%
1.06%±1.72%0.72%±2.50%0.76%±1.40%0.61%±1.36%0.15%±0.31%69,327
25
2231
Qwen3.7 Max
Alibaba · Proprietary
0.16%±0.88%
0.24%±2.11%5.07%±2.72%0.21%±1.55%4.20%±1.49%0.51%±0.26%27,375
26
2231
DeepSeek V4 Pro
DeepSeek · MIT
0.33%±0.86%
2.74%±2.16%3.35%±2.78%0.30%±1.54%3.82%±0.92%0.33%±0.27%28,084
27
2035
Kimi K2.6
Moonshot · Modified MIT
0.48%±2.16%
2.28%±4.11%2.12%±6.80%1.66%±4.49%6.27%±4.39%1.16%±0.16%10,970
28
2331
Gemini 3.1 Pro Preview
Google · Proprietary
0.54%±0.71%
1.71%±1.57%2.91%±2.23%2.78%±1.23%11.13%±1.34%1.01%±0.20%79,360
29
2331
Gemini 3.5 Flash (High)
Google · Proprietary
0.73%±0.59%
0.24%±1.40%1.09%±1.83%0.98%±1.05%1.98%±0.95%0.16%±0.20%90,015
30
2235
Tencent
Hy3
Tencent · Apache 2.0
1.12%±1.42%
2.20%±3.04%3.70%±4.91%9.01%±2.60%2.82%±1.82%0.89%±0.63%14,807
31
2436
Qwen3.7 Plus
Alibaba · Proprietary
1.87%±1.31%
0.07%±3.33%10.87%±3.94%3.78%±2.73%5.29%±1.96%0.07%±0.40%15,926
32
2936
Mimo V2.5 Pro
Xiaomi · MIT
2.42%±0.92%
3.88%±2.24%7.27%±2.82%2.19%±1.68%1.30%±1.45%0.08%±0.35%27,785
33
2936
Minimax M3
MiniMax · MiniMax Community License
2.47%±0.85%
5.15%±2.21%9.10%±2.65%4.19%±1.64%5.48%±0.82%0.59%±0.33%27,674
34
2936
DeepSeek V4 Flash
DeepSeek · MIT
2.70%±0.88%
3.12%±2.37%10.05%±2.68%1.90%±1.67%2.82%±0.97%1.25%±0.39%23,569
35
2936
gemini-3.6-flash
Google · Proprietary
2.73%±1.38%
1.01%±3.37%5.99%±4.30%4.56%±2.80%3.19%±1.82%1.11%±0.18%9,009
36
3138
Gemini 3.5 Flash (Medium)
Google · Proprietary
4.29%±1.49%
9.94%±3.70%5.61%±4.44%4.99%±2.87%1.19%±2.47%0.30%±0.82%11,957
37
3639
Thinking Machines
Inkling
Thinky · Apache 2.0
6.60%±0.89%
11.06%±2.48%16.01%±2.60%12.57%±1.77%6.15%±1.08%0.48%±0.24%32,583
38
3642
Mistral Medium 3.5
Mistral · Modified MIT
6.82%±2.16%
10.43%±5.02%8.71%±6.47%11.31%±4.31%0.13%±2.86%3.53%±2.85%4,844
39
3742
Grok 4.3 (High)
SpaceXAI · Proprietary
8.30%±0.83%
9.31%±1.74%13.95%±1.93%6.51%±1.26%12.72%±2.73%1.00%±0.16%60,015
40
3842
Grok Build 0.1
SpaceXAI · Proprietary
8.82%±0.87%
5.31%±1.80%11.12%±2.27%9.95%±1.54%18.59%±2.55%0.85%±0.16%71,454
41
3842
Gemini 3 Flash
Google · Proprietary
8.92%±0.76%
8.51%±1.64%11.36%±1.85%4.37%±1.20%20.36%±2.21%0.01%±0.77%80,624
42
3844
gemini-3.5-flash-lite
Google · Proprietary
10.37%±1.43%
14.10%±3.56%13.30%±3.78%10.72%±2.74%13.31%±3.09%0.45%±0.62%9,015
43
4244
Minimax M2.7
MiniMax · Modified MIT
11.59%±1.10%
13.21%±2.52%15.89%±2.95%13.41%±2.01%16.44%±2.91%1.01%±0.19%27,865
44
4246
Nemotron 3 Ultra
Nvidia · OpenMDW-1.1
13.75%±2.33%
16.20%±5.23%14.76%±6.61%19.08%±4.76%19.00%±5.52%0.30%±0.41%11,618
45
4446
Grok 4.3
SpaceXAI · Proprietary
14.86%±1.14%
11.77%±1.67%16.19%±1.82%6.44%±1.21%40.98%±4.91%1.06%±0.16%79,964
46
4446
Gemma 4 31B
Google · Apache 2.0
17.60%±2.38%
0.27%±1.93%2.90%±2.87%9.16%±1.83%44.31%±8.51%31.37%±7.31%56,459
Signal Leaders
  1. AnthropicClaude Opus 5 (Max)gets users to confirm the task is done most often16.36%±3.11%
  2. GPT 5.6 Sol (xHigh)draws the most positive responses relative to negative ones22.70%±6.42%
  3. AnthropicClaude Opus 5 (High)lands user corrections best10.29%±2.51%
  4. GPT 5.5 (xHigh)recovers from failed commands with the fewest steps14.13%±1.19%
  5. GPT 5.6 Luna (xHigh)least likely to hallucinate tools it doesn't have1.16%±0.16%

Confirmed Success

How often the model gets users to confirm the task is done.

  1. 1AnthropicClaude Opus 5 (Max)16.36%
    1AnthropicClaude Opus 5 (Max)16.36%
  2. 2AnthropicClaude Opus 5 (High)16.04%
    2AnthropicClaude Opus 5 (High)16.04%
  3. 3Kimi K3 (Max)14.97%
    3Kimi K3 (Max)14.97%
  4. 4AnthropicClaude Fable 5 (High)11.47%
    4AnthropicClaude Fable 5 (High)11.47%
  5. 5GLM 5.2 (Max)9.21%
    5GLM 5.2 (Max)9.21%
  6. 6GPT 5.6 Sol (xHigh)9.06%
    6GPT 5.6 Sol (xHigh)9.06%
  7. 7AnthropicClaude Opus 4.8 (Thinking)8.81%
    7AnthropicClaude Opus 4.8 (Thinking)8.81%
  8. 8AnthropicClaude Opus 4.88.78%
    8AnthropicClaude Opus 4.88.78%
  9. 9Deepseek V4 Flash (High) (20260731)8.53%
    9Deepseek V4 Flash (High) (20260731)8.53%
  10. 10AnthropicClaude Opus 4.7 (Thinking)6.65%
    10AnthropicClaude Opus 4.7 (Thinking)6.65%
813,263 Sessions

Praise vs Complaint

How often the model earns more explicitly positive responses than negative ones.

  1. 1GPT 5.6 Sol (xHigh)22.70%
    1GPT 5.6 Sol (xHigh)22.70%
  2. 2AnthropicClaude Fable 5 (High)22.56%
    2AnthropicClaude Fable 5 (High)22.56%
  3. 3AnthropicClaude Opus 4.8 (Thinking)21.46%
    3AnthropicClaude Opus 4.8 (Thinking)21.46%
  4. 4Kimi K3 (Max)19.68%
    4Kimi K3 (Max)19.68%
  5. 5AnthropicClaude Opus 5 (High)19.46%
    5AnthropicClaude Opus 5 (High)19.46%
  6. 6AnthropicClaude Opus 5 (Max)19.07%
    6AnthropicClaude Opus 5 (Max)19.07%
  7. 7AnthropicClaude Sonnet 5 (High)14.80%
    7AnthropicClaude Sonnet 5 (High)14.80%
  8. 8GPT 5.5 (xHigh)13.56%
    8GPT 5.5 (xHigh)13.56%
  9. 9AnthropicClaude Opus 4.813.52%
    9AnthropicClaude Opus 4.813.52%
  10. 10GLM 5.2 (Max)12.39%
    10GLM 5.2 (Max)12.39%
324,891 Sessions

Steerability

How well the model lands user corrections when they push back.

  1. 1AnthropicClaude Opus 5 (High)10.29%
    1AnthropicClaude Opus 5 (High)10.29%
  2. 2AnthropicClaude Fable 5 (High)10.01%
    2AnthropicClaude Fable 5 (High)10.01%
  3. 3AnthropicClaude Opus 4.79.54%
    3AnthropicClaude Opus 4.79.54%
  4. 4GPT 5.5 (High)9.09%
    4GPT 5.5 (High)9.09%
  5. 5AnthropicClaude Opus 4.7 (Thinking)8.61%
    5AnthropicClaude Opus 4.7 (Thinking)8.61%
  6. 6AnthropicClaude Opus 4.8 (Thinking)8.50%
    6AnthropicClaude Opus 4.8 (Thinking)8.50%
  7. 7GPT 5.6 Sol (xHigh)8.42%
    7GPT 5.6 Sol (xHigh)8.42%
  8. 8GPT 5.5 (xHigh)8.19%
    8GPT 5.5 (xHigh)8.19%
  9. 9AnthropicClaude Opus 4.88.10%
    9AnthropicClaude Opus 4.88.10%
  10. 10Kimi K3 (Max)7.45%
    10Kimi K3 (Max)7.45%
553,493 Sessions

Bash Recovery

How quickly the model recovers when a command doesn't work.

  1. 1GPT 5.5 (xHigh)14.13%
    1GPT 5.5 (xHigh)14.13%
  2. 2AnthropicClaude Opus 5 (Max)13.62%
    2AnthropicClaude Opus 5 (Max)13.62%
  3. 3AnthropicClaude Opus 4.7 (Thinking)13.47%
    3AnthropicClaude Opus 4.7 (Thinking)13.47%
  4. 4AnthropicClaude Fable 5 (High)13.10%
    4AnthropicClaude Fable 5 (High)13.10%
  5. 5AnthropicClaude Opus 5 (High)13.07%
    5AnthropicClaude Opus 5 (High)13.07%
  6. 6GPT 5.5 (High)12.32%
    6GPT 5.5 (High)12.32%
  7. 7Grok 4.511.57%
    7Grok 4.511.57%
  8. 8GPT 5.511.22%
    8GPT 5.511.22%
  9. 9GPT 5.6 Luna (xHigh)11.01%
    9GPT 5.6 Luna (xHigh)11.01%
  10. 10AnthropicClaude Opus 4.610.95%
    10AnthropicClaude Opus 4.610.95%
523,861 Sessions

Tool Hallucination

How much the model hallucinates tools it doesn't have.

  1. 1GPT 5.6 Luna (xHigh)1.16%
    1GPT 5.6 Luna (xHigh)1.16%
  2. 2GPT 5.6 Sol (xHigh)1.16%
    2GPT 5.6 Sol (xHigh)1.16%
  3. 3Kimi K2.61.16%
    3Kimi K2.61.16%
  4. 4AnthropicClaude Fable 5 (High)1.16%
    4AnthropicClaude Fable 5 (High)1.16%
  5. 5Kimi K2.7 Code1.16%
    5Kimi K2.7 Code1.16%
  6. 6GPT 5.6 Terra (xHigh)1.16%
    6GPT 5.6 Terra (xHigh)1.16%
  7. 7GPT 5.51.16%
    7GPT 5.51.16%
  8. 8Grok 4.51.16%
    8Grok 4.51.16%
  9. 9GPT 5.5 (High)1.16%
    9GPT 5.5 (High)1.16%
  10. 10Kimi K3 (Max)1.16%
    10Kimi K3 (Max)1.16%
1,804,218 Sessions

Frequently asked questions

Agent Mode

Try Agent Mode

Put these models to work on your own real tasks in Agent Mode.

Get started
How the Agent Leaderboard works

How the Agent Leaderboard works

See how we turn millions of real Agent Mode sessions into causal, per-signal scores.

Read the methodology

USE CASES

  • Chat with AI
  • Build Apps & Websites
  • Write & Edit Text
  • Search the Web
  • Generate Images
  • Generate Videos
  • Chose any model
  • Compare Models Side by Side

LEADERBOARD RANKINGS

  • Overall
  • Agent
  • Text
  • WebDev
  • Image-to-WebDev
  • Text to Image
  • Image Edit
  • Text to Video
  • Image to Video
  • Video Edit
  • Vision
  • Document
  • Search

COMPANY

  • About Us
  • How It Works
  • Blog
  • Careers
  • Changelog
  • Help Center
  • FAQ

LEGAL

  • Terms
  • Privacy
  • Cookies

FOLLOW

  • X
  • LinkedIn
  • YouTube
  • Discord

© Arena Intelligence 2026