View overall rankings across various AI models in text-to-text tasks across math, coding, creative writing, and other open-ended domains.
Lab Rank | Model Score | Rank Spread | ||
|---|---|---|---|---|
| 1 | Anthropic claude-fable-5 · Proprietary | 1543±26 | 1 | 16 |
| 2 | Google gemini-3.5-flash-high · Proprietary | 1522±25 | 2 | 115 |
| 3 | Alibaba qwen3.7-max-preview · Proprietary | 1495±40 | 7 | 157 |
| 4 | SpaceXAI grok-4.5 · Proprietary | 1494±34 | 8 | 150 |
| 5 | OpenAI gpt-5.4-high · Proprietary | 1490±11 | 9 | 335 |
| 6 | Baidu ernie-5.1 · Proprietary | 1481±14 | 15 | 544 |
| 7 | Z.ai glm-5.1 | 1477±15 | 16 | 548 |
| 8 | Xiaomi mimo-v2.5-pro | 1475±13 | 18 | 548 |
| 9 | Moonshot kimi-k2.6 | 1475±14 | 20 | 550 |
| 10 | Meta muse-spark-1.1 · Proprietary | 1468±33 | 27 | 398 |
| 11 | Thinky inkling | 1467±42 | 28 | 2114 |
| 12 | Nvidia nvidia-nemotron-3-ultra-550b-a55b-nvfp4 | 1460±25 | 33 | 698 |
| 13 | DeepSeek deepseek-v4-pro-thinking | 1458±13 | 36 | 1180 |
| 14 | MiniMax minimax-m3 | 1443±17 | 50 | 22114 |
| 15 | Meituan longcat-flash-chat | 1441±22 | 56 | 14119 |
| 16 | Amazon amazon-nova-experimental-chat-26-02-10 · Proprietary | 1440±39 | 57 | 7146 |
| 17 | Bytedance dola-seed-2.0-pro · Proprietary | 1438±10 | 61 | 29111 |
| 18 | Mistral mistral-medium-3.5 | 1436±25 | 64 | 16132 |
| 19 | Tencent hunyuan-hy3-preview | 1422±27 | 90 | 29152 |
| 20 | StepFun step-3.5-flash | 1408±11 | 114 | 69151 |
| 21 | intellect-3 | 1382±31 | 147 | 79184 |
| 22 | Ant Group ling-flash-2.0 | 1371±26 | 159 | 105188 |
| 23 | Ai2 olmo-3-32b-think | 1316±32 | 190 | 168225 |
| 24 | IBM granite-4.1-8b | 1315±39 | 192 | 161233 |
| 25 | Cohere command-a-03-2025 | 1299±9 | 205 | 187220 |
| 26 | Microsoft phi-4 | 1246±10 | 246 | 232262 |