View overall rankings across various AI models in text-to-text tasks across math, coding, creative writing, and other open-ended domains.
Lab Rank | Model Score | Rank Spread | ||
|---|---|---|---|---|
| 1 | Anthropic claude-fable-5 · Proprietary | 1507±7 | 1 | 15 |
| 2 | Meta muse-spark-1.1 · Proprietary | 1495±8 | 5 | 113 |
| 3 | Moonshot kimi-k3 · Proprietary | 1487±10 | 8 | 426 |
| 4 | OpenAI gpt-5.6-sol-xhigh · Proprietary | 1486±8 | 9 | 423 |
| 5 | Google gemini-3-pro · Proprietary | 1486±4 | 10 | 516 |
| 6 | Alibaba qwen3.7-max-preview · Proprietary | 1475±10 | 19 | 742 |
| 7 | SpaceXAI grok-4.20-beta1 · Proprietary | 1474±5 | 22 | 1136 |
| 8 | Z.ai glm-5.1 | 1470±5 | 29 | 1542 |
| 9 | Baidu ernie-5.1 · Proprietary | 1468±5 | 32 | 1642 |
| 10 | Xiaomi mimo-v2.5-pro | 1466±4 | 34 | 1845 |
| 11 | DeepSeek deepseek-v4-pro | 1457±4 | 45 | 3558 |
| 12 | Bytedance dola-seed-2.0-pro · Proprietary | 1456±4 | 48 | 3759 |
| 13 | Thinky inkling | 1447±9 | 61 | 4079 |
| 14 | MiniMax minimax-m3 | 1445±5 | 63 | 5276 |
| 15 | Meituan longcat-flash-chat-2602-exp · Proprietary | 1436±5 | 76 | 6290 |
| 16 | Mistral mistral-medium-3.5 | 1428±7 | 87 | 71111 |
| 17 | Amazon amazon-nova-experimental-chat-26-02-10 · Proprietary | 1427±10 | 88 | 69114 |
| 18 | Nvidia nvidia-nemotron-3-ultra-550b-a55b-nvfp4 | 1424±7 | 94 | 74114 |
| 19 | Tencent hunyuan-hy3-preview | 1412±8 | 116 | 93131 |
| 20 | StepFun step-3.5-flash | 1395±4 | 138 | 126150 |
| 21 | trinity-large-preview | 1379±4 | 156 | 146164 |
| 22 | Cohere command-a-03-2025 | 1354±3 | 178 | 171197 |
| 23 | Inception AI mercury-2 · Proprietary | 1347±11 | 190 | 171211 |
| 24 | Ant Group ling-flash-2.0 | 1346±7 | 191 | 175209 |
| 25 | Ai2 olmo-3.1-32b-instruct | 1330±6 | 210 | 195226 |
| 26 | IBM granite-4.1-8b | 1307±10 | 241 | 219252 |
| 27 | Microsoft phi-4 | 1256±5 | 282 | 276284 |