• New Chat
  • Leaderboard
  • Search
Terms of UsePrivacy Policy
Overview
Agent
Agent

Agent ArenaView Methodology

Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.

Aug 13, 2026
1,793,983 sessions
48 models
Model
1
16
Anthropic
Claude Opus 5 (High)
Anthropic · Proprietary
12.19%±1.45%
15.35%±3.03%19.26%±5.29%11.34%±2.79%13.89%±0.79%1.13%±0.19%19,739$5 / $25
2
18
Anthropic
Claude Fable 5 (High)
Anthropic · Proprietary
12.01%±2.57%
10.66%±4.90%25.14%±9.01%8.84%±5.23%14.19%±3.71%1.22%±0.19%24,417$10 / $50
3
16
Anthropic
Claude Opus 5 (Max)
Anthropic · Proprietary
11.95%±1.71%
17.96%±3.18%19.30%±6.26%6.95%±3.43%14.38%±0.88%1.18%±0.19%15,515$5 / $25
4
111
GPT 5.6 Sol (xHigh)
OpenAI · Proprietary
10.86%±1.80%
10.12%±3.69%23.03%±6.72%9.23%±3.74%10.71%±1.27%1.22%±0.19%18,091$5 / $30
5
18
Kimi K3 (Max)
Moonshot · Kimi K3 license
10.60%±1.04%
15.55%±2.06%20.31%±3.78%7.80%±1.91%8.12%±0.83%1.22%±0.19%28,369$3 / $15
6
113
Anthropic
Claude Opus 4.8 (High)
Anthropic · Proprietary
9.78%±1.78%
9.45%±3.10%22.52%±5.71%8.44%±3.33%9.29%±2.89%0.80%±2.47%35,151$5 / $25
7
313
GPT 5.5 (xHigh)
OpenAI · Proprietary
8.90%±1.03%
4.97%±2.18%14.61%±3.62%9.17%±1.92%14.53%±1.25%1.21%±0.19%47,586$2.50 / $15
8
316
Anthropic
Claude Opus 4.7 (High)
Anthropic · Proprietary
8.17%±1.43%
6.61%±2.95%12.32%±4.81%7.72%±2.91%13.09%±2.08%1.12%±0.21%36,116$5 / $25
9
616
GPT 5.5 (High)
OpenAI · Proprietary
7.73%±0.97%
4.03%±2.04%11.62%±3.39%8.59%±1.79%13.21%±0.97%1.22%±0.19%72,849$2.50 / $15
10
516
Anthropic
Claude Opus 4.7
Anthropic · Proprietary
7.66%±1.45%
5.45%±3.08%11.57%±4.71%9.95%±2.87%10.15%±2.31%1.17%±0.19%36,677$5 / $25
12
520
Anthropic
Claude Sonnet 5 (High)
Anthropic · Proprietary
7.14%±2.29%
3.13%±4.73%15.08%±7.93%5.22%±4.88%11.21%±1.80%1.07%±0.20%25,733$1 / $5
13
618
Anthropic
Claude Opus 4.6
Anthropic · Proprietary
6.89%±1.39%
5.28%±2.98%8.21%±4.53%8.30%±2.73%11.46%±1.80%1.22%±0.19%35,853$5 / $25
14
818
GLM 5.2 (Max)
Z.ai · MIT · SiliconFlow
6.74%±0.90%
8.39%±1.91%11.60%±3.18%6.14%±1.58%6.33%±1.07%1.22%±0.19%53,469$1.40 / $4.40
15
818
GPT 5.5
OpenAI · Proprietary
6.35%±0.91%
3.63%±1.98%7.71%±3.12%7.36%±1.72%11.86%±1.00%1.22%±0.19%73,997$2.50 / $15
16
819
Grok 4.5
SpaceXAI · Proprietary
6.19%±1.20%
6.42%±2.68%5.04%±4.27%7.04%±2.29%11.22%±1.20%1.22%±0.19%29,943$2 / $6
17
1223
GPT 5.4 (High)
OpenAI · Proprietary
4.99%±0.95%
4.33%±2.06%3.82%±3.18%6.05%±1.81%9.55%±1.32%1.22%±0.19%73,220$2.50 / $15
18
1124
GPT 5.6 Luna (xHigh)
OpenAI · Proprietary
4.28%±1.88%
1.09%±4.27%8.16%±6.50%1.48%±3.85%11.65%±1.45%1.22%±0.19%8,692$1 / $6
19
1624
Deepseek V4 Flash (High) (20260731)
DeepSeek · MIT
4.04%±0.81%
8.70%±1.93%2.46%±2.71%3.34%±1.57%4.48%±0.69%1.21%±0.19%36,264N/A
20
1524
GPT 5.6 Terra (xHigh)
OpenAI · Proprietary
3.78%±1.25%
1.80%±3.04%3.56%±4.22%6.22%±2.43%9.70%±1.15%1.22%±0.19%14,022$2.50 / $15
21
1724
Gemini 3.7 Flash (High)
Google · Proprietary
3.61%±1.22%
9.85%±2.83%1.84%±4.04%2.83%±2.51%2.33%±1.39%1.18%±0.19%13,000$0.75 / $3.57
22
1726
Anthropic
Claude Sonnet 4.6
Anthropic · Proprietary
3.12%±1.40%
0.06%±3.15%1.07%±4.17%2.21%±2.70%11.15%±2.72%1.09%±0.24%36,710$1.50 / $7.50
23
1732
Anthropic
Claude Opus 4.8
Anthropic · Proprietary
2.12%±2.72%
8.76%±3.17%13.67%±5.40%8.34%±3.22%10.86%±2.23%31.01%±10.90%33,192$5 / $25
24
2229
Meta
Muse Spark 1.1
Meta · Proprietary
1.17%±0.61%
7.28%±1.40%4.95%±1.82%3.25%±1.15%5.59%±1.20%1.19%±0.19%72,668$1.25 / $4.25
25
1834
Kimi K2.7 Code
Moonshot · Modified MIT
1.08%±2.15%
4.43%±4.55%2.74%±7.27%1.76%±4.86%1.22%±2.58%1.22%±0.19%11,029$0.47 / $2
26
2332
GLM 5.1
Z.ai · MIT · SiliconFlow
0.58%±0.82%
1.75%±1.82%0.84%±2.56%1.73%±1.48%0.91%±1.47%0.53%±0.37%71,334$1.40 / $4.40
27
2333
Qwen3.7 Max
Alibaba · Proprietary
0.20%±0.87%
0.34%±2.13%4.24%±2.67%0.03%±1.52%5.02%±1.47%0.60%±0.27%30,743N/A
28
2333
DeepSeek V4 Pro
DeepSeek · MIT
0.14%±0.88%
2.16%±2.21%2.50%±2.82%0.59%±1.63%4.41%±0.98%0.35%±0.29%30,083$1.74 / $3.48
29
2433
Gemini 3.5 Flash (High)
Google · Proprietary
0.29%±0.64%
0.63%±1.50%0.28%±1.99%0.47%±1.17%1.66%±1.00%0.32%±0.22%93,770$0.75 / $4.50
30
2435
Gemini 3.1 Pro Preview
Google · Proprietary
0.44%±0.75%
1.49%±1.69%3.65%±2.32%3.21%±1.32%11.48%±1.43%0.93%±0.31%81,349$1 / $6
31
2238
Kimi K2.6
Moonshot · Modified MIT
0.56%±2.31%
0.56%±4.67%2.05%±7.32%1.15%±4.96%5.46%±4.40%1.22%±0.19%11,177$0.95 / $4
32
2438
Tencent
Hy3
Tencent · Apache 2.0
1.31%±1.27%
2.47%±2.84%1.80%±4.39%8.63%±2.38%4.01%±1.56%1.27%±0.71%18,388$0.13 / $0.53
33
2938
Mimo V2.5 Pro
Xiaomi · MIT
1.97%±0.91%
2.81%±2.21%6.62%±2.77%1.98%±1.63%1.57%±1.59%0.01%±0.34%31,435$0.43 / $0.87
34
2638
Qwen3.7 Plus
Alibaba · Proprietary
1.98%±1.34%
1.10%±3.43%9.74%±4.03%5.57%±2.78%6.35%±1.81%0.17%±0.40%16,766$0.32 / $1.28
35
3138
DeepSeek V4 Flash
DeepSeek · MIT
2.20%±0.90%
2.02%±2.43%9.24%±2.73%1.00%±1.69%2.61%±1.02%1.33%±0.42%23,570$0.14 / $0.28
36
3038
Gemini 3.6 Flash (High)
Google · Proprietary
2.36%±1.22%
0.87%±2.99%4.12%±3.90%4.42%±2.38%3.57%±1.69%1.18%±0.20%12,411$0.38 / $1.88
37
3138
Minimax M3
MiniMax · MiniMax Community License
2.53%±0.84%
5.81%±2.22%8.24%±2.60%5.14%±1.58%5.87%±0.83%0.69%±0.33%31,143$0.60 / $2.40
38
3138
Gemini 3.5 Flash (Medium)
Google · Proprietary
3.48%±1.43%
8.34%±3.61%5.74%±4.26%3.50%±2.82%0.34%±2.26%0.51%±0.67%12,965$0.75 / $4.50
39
3940
Thinking Machines
Inkling
Thinky · Apache 2.0
6.57%±0.91%
11.70%±2.59%16.70%±2.61%11.78%±1.86%6.79%±1.08%0.56%±0.27%35,781$0.50 / $2.02
40
3944
Mistral Medium 3.5
Mistral · Modified MIT
7.03%±2.00%
9.72%±4.62%9.83%±5.70%11.01%±4.03%1.69%±3.56%2.87%±2.12%5,730$1.50 / $7.50
41
4045
Grok 4.3 (High)
SpaceXAI · Proprietary
8.47%±0.89%
9.71%±1.86%13.27%±2.02%7.20%±1.37%13.25%±2.98%1.09%±0.19%62,174$1.25 / $2.50
42
4045
Gemini 3 Flash
Google · Proprietary
8.50%±0.87%
7.16%±1.70%10.22%±1.95%3.43%±1.29%21.02%±2.39%0.65%±1.86%82,748$0.50 / $3
43
4045
Grok Build 0.1
SpaceXAI · Proprietary
9.10%±0.95%
5.37%±1.92%11.02%±2.32%8.61%±1.68%21.44%±2.91%0.96%±0.19%73,718N/A
44
4047
Solar Pro 4
Upstage · Proprietary
10.10%±2.23%
8.44%±5.67%15.67%±6.83%13.32%±4.02%13.56%±4.52%0.51%±0.53%4,851$0.03 / $0.12
45
4146
Gemini 3.5 Flash Lite
Google · Proprietary
10.38%±1.19%
12.73%±2.91%13.51%±3.03%10.28%±2.25%15.08%±2.74%0.28%±0.50%20,237$0.15 / $1.25
46
4447
Minimax M2.7
MiniMax · Modified MIT
11.17%±1.08%
10.88%±2.51%15.51%±2.87%13.67%±1.96%16.82%±2.89%1.03%±0.23%31,290$0.30 / $1.20
47
4549
Nemotron 3 Ultra
Nvidia · OpenMDW-1.1
14.63%±2.54%
15.38%±5.41%14.33%±6.73%19.91%±4.74%24.09%±7.19%0.54%±0.35%11,997N/A
48
4748
Grok 4.3
SpaceXAI · Proprietary
14.67%±1.27%
10.83%±1.76%16.23%±1.84%5.65%±1.30%41.77%±5.52%1.14%±0.19%82,146$1.25 / $2.50
49
4849
Gemma 4 31B
Google · Apache 2.0
18.88%±2.85%
0.26%±2.07%2.58%±3.00%9.36%±2.05%51.50%±11.13%31.21%±7.86%56,561$0.14 / $0.40
Signal Leaders
  1. AnthropicClaude Opus 5 (Max)gets users to confirm the task is done most often17.96%±3.18%
  2. AnthropicClaude Fable 5 (High)draws the most positive responses relative to negative ones25.14%±9.01%
  3. AnthropicClaude Opus 5 (High)lands user corrections best11.34%±2.79%
  4. GPT 5.5 (xHigh)recovers from failed commands with the fewest steps14.53%±1.25%
  5. GPT 5.6 Terra (xHigh)least likely to hallucinate tools it doesn't have1.22%±0.19%

Confirmed Success

How often the model gets users to confirm the task is done.

  1. 1AnthropicClaude Opus 5 (Max)17.96%
    1AnthropicClaude Opus 5 (Max)17.96%
  2. 2Kimi K3 (Max)15.55%
    2Kimi K3 (Max)15.55%
  3. 3AnthropicClaude Opus 5 (High)15.35%
    3AnthropicClaude Opus 5 (High)15.35%
  4. 4AnthropicClaude Fable 5 (High)10.66%
    4AnthropicClaude Fable 5 (High)10.66%
  5. 5GPT 5.6 Sol (xHigh)10.12%
    5GPT 5.6 Sol (xHigh)10.12%
  6. 6Gemini 3.7 Flash (High)9.85%
    6Gemini 3.7 Flash (High)9.85%
  7. 7AnthropicClaude Opus 4.8 (High)9.45%
    7AnthropicClaude Opus 4.8 (High)9.45%
  8. 8AnthropicClaude Opus 4.88.76%
    8AnthropicClaude Opus 4.88.76%
  9. 9Deepseek V4 Flash (High) (20260731)8.70%
    9Deepseek V4 Flash (High) (20260731)8.70%
  10. 10GLM 5.2 (Max)8.39%
    10GLM 5.2 (Max)8.39%
905,643 Sessions

Praise vs Complaint

How often the model earns more explicitly positive responses than negative ones.

  1. 1AnthropicClaude Fable 5 (High)25.14%
    1AnthropicClaude Fable 5 (High)25.14%
  2. 2GPT 5.6 Sol (xHigh)23.03%
    2GPT 5.6 Sol (xHigh)23.03%
  3. 3AnthropicClaude Opus 4.8 (High)22.52%
    3AnthropicClaude Opus 4.8 (High)22.52%
  4. 4Kimi K3 (Max)20.31%
    4Kimi K3 (Max)20.31%
  5. 5AnthropicClaude Opus 5 (Max)19.30%
    5AnthropicClaude Opus 5 (Max)19.30%
  6. 6AnthropicClaude Opus 5 (High)19.26%
    6AnthropicClaude Opus 5 (High)19.26%
  7. 7AnthropicClaude Sonnet 5 (High)15.08%
    7AnthropicClaude Sonnet 5 (High)15.08%
  8. 8GPT 5.5 (xHigh)14.61%
    8GPT 5.5 (xHigh)14.61%
  9. 9AnthropicClaude Opus 4.813.67%
    9AnthropicClaude Opus 4.813.67%
  10. 10AnthropicClaude Opus 4.7 (High)12.32%
    10AnthropicClaude Opus 4.7 (High)12.32%
366,933 Sessions

Steerability

How well the model lands user corrections when they push back.

  1. 1AnthropicClaude Opus 5 (High)11.34%
    1AnthropicClaude Opus 5 (High)11.34%
  2. 2AnthropicClaude Opus 4.79.95%
    2AnthropicClaude Opus 4.79.95%
  3. 3GPT 5.6 Sol (xHigh)9.23%
    3GPT 5.6 Sol (xHigh)9.23%
  4. 4GPT 5.5 (xHigh)9.17%
    4GPT 5.5 (xHigh)9.17%
  5. 5AnthropicClaude Fable 5 (High)8.84%
    5AnthropicClaude Fable 5 (High)8.84%
  6. 6GPT 5.5 (High)8.59%
    6GPT 5.5 (High)8.59%
  7. 7AnthropicClaude Opus 4.8 (High)8.44%
    7AnthropicClaude Opus 4.8 (High)8.44%
  8. 8AnthropicClaude Opus 4.88.34%
    8AnthropicClaude Opus 4.88.34%
  9. 9AnthropicClaude Opus 4.68.30%
    9AnthropicClaude Opus 4.68.30%
  10. 10Kimi K3 (Max)7.80%
    10Kimi K3 (Max)7.80%
612,356 Sessions

Bash Recovery

How quickly the model recovers when a command doesn't work.

  1. 1GPT 5.5 (xHigh)14.53%
    1GPT 5.5 (xHigh)14.53%
  2. 2AnthropicClaude Opus 5 (Max)14.38%
    2AnthropicClaude Opus 5 (Max)14.38%
  3. 3AnthropicClaude Fable 5 (High)14.19%
    3AnthropicClaude Fable 5 (High)14.19%
  4. 4AnthropicClaude Opus 5 (High)13.89%
    4AnthropicClaude Opus 5 (High)13.89%
  5. 5GPT 5.5 (High)13.21%
    5GPT 5.5 (High)13.21%
  6. 6AnthropicClaude Opus 4.7 (High)13.09%
    6AnthropicClaude Opus 4.7 (High)13.09%
  7. 7GPT 5.511.86%
    7GPT 5.511.86%
  8. 8GPT 5.6 Luna (xHigh)11.65%
    8GPT 5.6 Luna (xHigh)11.65%
  9. 9AnthropicClaude Opus 4.611.46%
    9AnthropicClaude Opus 4.611.46%
  10. 10Grok 4.511.22%
    10Grok 4.511.22%
607,862 Sessions

Tool Hallucination

How much the model hallucinates tools it doesn't have.

  1. 1GPT 5.6 Terra (xHigh)1.22%
    1GPT 5.6 Terra (xHigh)1.22%
  2. 2Kimi K3 (Max)1.22%
    2Kimi K3 (Max)1.22%
  3. 3GPT 5.51.22%
    3GPT 5.51.22%
  4. 4GPT 5.4 (High)1.22%
    4GPT 5.4 (High)1.22%
  5. 5Grok 4.51.22%
    5Grok 4.51.22%
  6. 6GLM 5.2 (Max)1.22%
    6GLM 5.2 (Max)1.22%
  7. 7GPT 5.6 Sol (xHigh)1.22%
    7GPT 5.6 Sol (xHigh)1.22%
  8. 8GPT 5.5 (High)1.22%
    8GPT 5.5 (High)1.22%
  9. 9GPT 5.6 Luna (xHigh)1.22%
    9GPT 5.6 Luna (xHigh)1.22%
  10. 10Kimi K2.61.22%
    10Kimi K2.61.22%
2,030,546 Sessions

Frequently asked questions

Agent Mode

Try Agent Mode

Put these models to work on your own real tasks in Agent Mode.

Get started
How the Agent Leaderboard works

How the Agent Leaderboard works

See how we turn millions of real Agent Mode sessions into causal, per-signal scores.

Read the methodology

USE CASES

  • Chat with AI
  • Build Apps & Websites
  • Write & Edit Text
  • Search the Web
  • Generate Images
  • Generate Videos
  • Chose any model
  • Compare Models Side by Side

LEADERBOARD RANKINGS

  • Overall
  • Agent
  • Text
  • WebDev
  • Image-to-WebDev
  • Text to Image
  • Image Edit
  • Text to Video
  • Image to Video
  • Video Edit
  • Vision
  • Document
  • Search

COMPANY

  • About Us
  • How It Works
  • Blog
  • Careers
  • Changelog
  • Help Center
  • FAQ

LEGAL

  • Terms
  • Privacy
  • Cookies

FOLLOW

  • X
  • LinkedIn
  • YouTube
  • Discord

© Arena Intelligence 2026