• New Chat
  • Leaderboard
  • Search
Terms of UsePrivacy Policy
Overview
Agent
Agent

Agent ArenaView Methodology

Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.

Jul 28, 2026
1,412,751 sessions
14 open source models
Model
5
110
Kimi K3 (Max)
Moonshot · Kimi K3 license
9.91%±1.19%
14.23%±2.46%17.45%±4.26%9.67%±2.19%6.86%±1.03%1.32%±0.19%19,586
12
615
GLM 5.2 (Max)
Z.ai · MIT · SiliconFlow
7.08%±0.96%
9.08%±1.94%13.79%±3.42%5.96%±1.69%5.23%±1.09%1.32%±0.19%43,277
22
2130
GLM 5.1
Z.ai · MIT · SiliconFlow
0.44%±0.79%
1.04%±1.76%0.54%±2.60%0.15%±1.45%0.28%±1.31%0.50%±0.33%62,635
23
1731
Kimi K2.7 Code
Moonshot · Modified MIT
0.44%±1.82%
4.73%±3.63%3.42%±6.42%3.71%±3.65%3.56%±2.87%1.32%±0.19%10,359
26
2133
DeepSeek V4 Pro
DeepSeek · MIT
0.78%±0.94%
3.61%±2.41%5.37%±3.02%0.39%±1.74%4.77%±0.98%0.72%±0.26%21,459
27
2134
Tencent
Hy3
Tencent · Apache 2.0
1.03%±1.88%
1.55%±3.99%2.83%±6.51%8.17%±3.49%1.44%±2.58%0.33%±0.52%8,506
29
2134
Kimi K2.6
Moonshot · Modified MIT
1.09%±1.95%
3.10%±3.68%1.58%±6.21%4.90%±3.76%6.56%±3.93%1.32%±0.19%10,436
32
2634
Mimo V2.5 Pro
Xiaomi · MIT
2.52%±1.01%
3.81%±2.46%8.05%±3.13%2.11%±1.89%1.09%±1.64%0.30%±0.37%21,399
33
2634
Minimax M3
MiniMax · MiniMax Community License
2.55%±0.93%
5.67%±2.43%9.01%±2.95%4.56%±1.83%5.71%±0.93%0.77%±0.37%20,996
34
2835
DeepSeek V4 Flash
DeepSeek · MIT
2.96%±0.95%
4.90%±2.51%8.38%±2.97%3.98%±1.78%3.40%±0.98%0.92%±0.42%20,775
36
3536
Thinking Machines
Inkling
Thinky · Apache 2.0
5.71%±0.87%
7.07%±2.35%15.75%±2.56%12.07%±1.80%5.92%±1.03%0.41%±0.28%25,699
41
4042
Minimax M2.7
MiniMax · Modified MIT
11.59%±1.17%
14.67%±2.82%15.74%±3.27%14.94%±2.22%13.69%±2.81%1.08%±0.26%21,221
42
4044
Nemotron 3 Ultra
Nvidia · OpenMDW-1.1
13.14%±2.33%
15.00%±5.18%13.58%±6.94%20.73%±4.93%16.58%±5.20%0.19%±0.53%10,820
44
4244
Gemma 4 31B
Google · Apache 2.0
16.10%±1.94%
1.00%±1.81%3.90%±2.70%8.47%±1.62%38.46%±6.49%28.69%±6.28%55,900
Signal Leaders
  1. Kimi K3 (Max)gets users to confirm the task is done most often14.23%±2.46%
  2. Kimi K3 (Max)draws the most positive responses relative to negative ones17.45%±4.26%
  3. Kimi K3 (Max)lands user corrections best9.67%±2.19%
  4. Kimi K3 (Max)recovers from failed commands with the fewest steps6.86%±1.03%
  5. Kimi K3 (Max)least likely to hallucinate tools it doesn't have1.32%±0.19%

Confirmed Success

How often the model gets users to confirm the task is done.

  1. 1Kimi K3 (Max)14.23%
    1Kimi K3 (Max)14.23%
  2. 2GLM 5.2 (Max)9.08%
    2GLM 5.2 (Max)9.08%
  3. 3Kimi K2.7 Code4.73%
    3Kimi K2.7 Code4.73%
  4. 4Kimi K2.63.10%
    4Kimi K2.63.10%
  5. 5GLM 5.11.04%
    5GLM 5.11.04%
  6. 6Gemma 4 31B1.00%
    6Gemma 4 31B1.00%
  7. 7TencentHy31.55%
    7TencentHy31.55%
  8. 8DeepSeek V4 Pro3.61%
    8DeepSeek V4 Pro3.61%
  9. 9Mimo V2.5 Pro3.81%
    9Mimo V2.5 Pro3.81%
  10. 10DeepSeek V4 Flash4.90%
    10DeepSeek V4 Flash4.90%
665,756 Sessions

Praise vs Complaint

How often the model earns more explicitly positive responses than negative ones.

  1. 1Kimi K3 (Max)17.45%
    1Kimi K3 (Max)17.45%
  2. 2GLM 5.2 (Max)13.79%
    2GLM 5.2 (Max)13.79%
  3. 3Kimi K2.7 Code3.42%
    3Kimi K2.7 Code3.42%
  4. 4TencentHy32.83%
    4TencentHy32.83%
  5. 5Kimi K2.61.58%
    5Kimi K2.61.58%
  6. 6GLM 5.10.54%
    6GLM 5.10.54%
  7. 7Gemma 4 31B3.90%
    7Gemma 4 31B3.90%
  8. 8DeepSeek V4 Pro5.37%
    8DeepSeek V4 Pro5.37%
  9. 9Mimo V2.5 Pro8.05%
    9Mimo V2.5 Pro8.05%
  10. 10DeepSeek V4 Flash8.38%
    10DeepSeek V4 Flash8.38%
265,857 Sessions

Steerability

How well the model lands user corrections when they push back.

  1. 1Kimi K3 (Max)9.67%
    1Kimi K3 (Max)9.67%
  2. 2GLM 5.2 (Max)5.96%
    2GLM 5.2 (Max)5.96%
  3. 3GLM 5.10.15%
    3GLM 5.10.15%
  4. 4DeepSeek V4 Pro0.39%
    4DeepSeek V4 Pro0.39%
  5. 5Mimo V2.5 Pro2.11%
    5Mimo V2.5 Pro2.11%
  6. 6Kimi K2.7 Code3.71%
    6Kimi K2.7 Code3.71%
  7. 7DeepSeek V4 Flash3.98%
    7DeepSeek V4 Flash3.98%
  8. 8Minimax M34.56%
    8Minimax M34.56%
  9. 9Kimi K2.64.90%
    9Kimi K2.64.90%
  10. 10TencentHy38.17%
    10TencentHy38.17%
441,879 Sessions

Bash Recovery

How quickly the model recovers when a command doesn't work.

  1. 1Kimi K3 (Max)6.86%
    1Kimi K3 (Max)6.86%
  2. 2Thinking MachinesInkling5.92%
    2Thinking MachinesInkling5.92%
  3. 3Minimax M35.71%
    3Minimax M35.71%
  4. 4GLM 5.2 (Max)5.23%
    4GLM 5.2 (Max)5.23%
  5. 5DeepSeek V4 Pro4.77%
    5DeepSeek V4 Pro4.77%
  6. 6DeepSeek V4 Flash3.40%
    6DeepSeek V4 Flash3.40%
  7. 7TencentHy31.44%
    7TencentHy31.44%
  8. 8Mimo V2.5 Pro1.09%
    8Mimo V2.5 Pro1.09%
  9. 9GLM 5.10.28%
    9GLM 5.10.28%
  10. 10Kimi K2.7 Code3.56%
    10Kimi K2.7 Code3.56%
415,292 Sessions

Tool Hallucination

How much the model hallucinates tools it doesn't have.

  1. 1Kimi K3 (Max)1.32%
    1Kimi K3 (Max)1.32%
  2. 2GLM 5.2 (Max)1.32%
    2GLM 5.2 (Max)1.32%
  3. 3Kimi K2.7 Code1.32%
    3Kimi K2.7 Code1.32%
  4. 4Kimi K2.61.32%
    4Kimi K2.61.32%
  5. 5Minimax M2.71.08%
    5Minimax M2.71.08%
  6. 6Minimax M30.77%
    6Minimax M30.77%
  7. 7DeepSeek V4 Pro0.72%
    7DeepSeek V4 Pro0.72%
  8. 8GLM 5.10.50%
    8GLM 5.10.50%
  9. 9Thinking MachinesInkling0.41%
    9Thinking MachinesInkling0.41%
  10. 10TencentHy30.33%
    10TencentHy30.33%
1,489,963 Sessions

Frequently asked questions

Agent Mode

Try Agent Mode

Put these models to work on your own real tasks in Agent Mode.

Get started
How the Agent Leaderboard works

How the Agent Leaderboard works

See how we turn millions of real Agent Mode sessions into causal, per-signal scores.

Read the methodology

USE CASES

  • Chat with AI
  • Build Apps & Websites
  • Write & Edit Text
  • Search the Web
  • Generate Images
  • Generate Videos
  • Chose any model
  • Compare Models Side by Side

LEADERBOARD RANKINGS

  • Overall
  • Agent
  • Text
  • WebDev
  • Image-to-WebDev
  • Text to Image
  • Image Edit
  • Text to Video
  • Image to Video
  • Video Edit
  • Vision
  • Document
  • Search

COMPANY

  • About Us
  • How It Works
  • Blog
  • Careers
  • Changelog
  • Help Center
  • FAQ

LEGAL

  • Terms
  • Privacy
  • Cookies

FOLLOW

  • X
  • LinkedIn
  • YouTube
  • Discord

© Arena Intelligence 2026