LLM performance leaderboard
Price, output speed and quality across every major model, compared side by side. This is the table we open before deciding what a workflow should run on. The data is published by Artificial Analysis — we host it here for convenience and credit it in full.
| Quality index | Cost $/1M tokens | Speed | Benchmarks | ||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| # | Model | ||||||||||||||||||||||||
| 1 | Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic | Jul 2026 | 63.1 | 78.0 | — | $10.00 | $5.00 | $25.00 | 54 | 34.57 | 34.57 | 93.2% | 54.9% | 55.7% | 75.7% | — | — | — | — | — | — | — | 89.1% | — | 42.1% |
| 2 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Anthropic | Jul 2026 | 62.5 | 77.0 | — | $10.00 | $5.00 | $25.00 | 54 | 18.21 | 18.21 | 93.7% | 54.4% | 55.0% | 76.3% | — | — | — | — | — | — | — | 88.0% | — | 43.3% |
| 3 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Anthropic | Jun 2026 | 62.1 | 76.5 | — | $20.00 | $10.00 | $50.00 | 62 | 54.13 | 54.13 | 92.6% | 55.5% | 60.2% | 76.7% | 63.5% | 98.5% | 62.9% | — | — | — | — | 84.6% | — | 38.1% |
| 4 | Claude Opus 5 (Adaptive Reasoning, High Effort) Anthropic | Jul 2026 | 61.5 | 76.5 | — | $10.00 | $5.00 | $25.00 | 53 | 11.33 | 11.33 | 93.7% | 52.8% | 54.3% | 76.3% | — | — | — | — | — | — | — | 87.6% | — | 44.7% |
| 5 | GPT-5.6 Sol (max) OpenAI | Jul 2026 | 60.9 | 77.4 | — | $8.00 | $4.00 | $20.00 | 84 | 50.56 | 50.56 | 94.1% | 49.5% | 56.1% | 77.7% | 72.7% | 85.1% | 65.9% | — | — | — | — | 88.0% | — | 44.3% |
| 6 | Grok 4.6 (high) SpaceXAI | Aug 2026 | 60.9 | 76.8 | — | $3.00 | $2.00 | $6.00 | 58 | 36.59 | 36.59 | 94.9% | 42.9% | 53.6% | 75.0% | — | — | — | — | — | — | — | 88.4% | — | 50.7% |
| 7 | Grok 4.6 (xhigh) SpaceXAI | Aug 2026 | 60.0 | 75.9 | — | $3.00 | $2.00 | $6.00 | 60 | 38.68 | 38.68 | 93.5% | 44.1% | 51.6% | 75.7% | — | — | — | — | — | — | — | 88.0% | — | 43.3% |
| 8 | Kimi K3 (max) Kimi | Jul 2026 | 59.7 | 76.2 | — | $6.00 | $3.00 | $15.00 | 36 | 3.56 | 58.53 | 93.5% | 46.9% | 58.7% | 82.7% | — | — | — | — | — | — | — | 85.0% | — | 46.0% |
| 9 | GLM-5.3 (max) Z AI | Aug 2026 | 59.5 | 74.8 | — | $2.15 | $1.40 | $4.40 | 75 | 1.61 | 28.20 | 91.7% | 42.3% | 56.5% | 76.3% | — | — | — | — | — | — | — | 83.9% | — | 50.3% |
| 10 | GPT-5.6 Sol (xhigh) OpenAI | Jul 2026 | 59.0 | 78.3 | — | $8.00 | $4.00 | $20.00 | 82 | 23.49 | 23.49 | 93.1% | 47.3% | 56.0% | 76.3% | 71.0% | 84.8% | 61.4% | — | — | — | — | 89.5% | — | 38.1% |
| 11 | Grok 4.6 (medium) SpaceXAI | Aug 2026 | 59.0 | 74.4 | — | $3.00 | $2.00 | $6.00 | 57 | 31.70 | 31.70 | 93.5% | 42.1% | 54.6% | 72.7% | — | — | — | — | — | — | — | 84.3% | — | 44.3% |
| 12 | Claude Opus 5 (Adaptive Reasoning, Medium Effort) Anthropic | Jul 2026 | 58.6 | 74.3 | — | $10.00 | $5.00 | $25.00 | 51 | 8.65 | 8.65 | 91.9% | 51.3% | 50.7% | 78.7% | — | — | — | — | — | — | — | 86.1% | — | 38.6% |
| 13 | Qwen3.8 Max Alibaba | Aug 2026 | 58.1 | 71.8 | — | $3.00 | $2.00 | $6.00 | 30 | 1.76 | 68.67 | 92.7% | 43.0% | 52.9% | 74.3% | — | — | — | — | — | — | — | 81.3% | — | 51.3% |
| 14 | Qwen3.8 2.4T A95B Alibaba | Aug 2026 | 57.7 | 71.9 | — | $3.00 | $2.00 | $6.00 | 31 | 1.91 | 65.63 | 93.5% | 42.4% | 51.6% | 75.3% | — | — | — | — | — | — | — | 82.0% | — | 49.1% |
| 15 | GLM-5.3-Flash Z AI | Aug 2026 | 57.5 | 71.5 | — | $0.24 | $0.15 | $0.50 | 50 | 1.11 | 41.02 | 91.2% | 39.9% | 46.1% | 78.0% | — | — | — | — | — | — | — | 84.3% | — | 47.2% |
| 16 | GPT-5.6 Sol (high) OpenAI | Jul 2026 | 57.3 | 77.2 | — | $8.00 | $4.00 | $20.00 | 78 | 14.11 | 14.11 | 92.8% | 46.0% | 56.9% | 75.3% | 69.2% | 83.3% | 62.1% | — | — | — | — | 87.3% | — | 36.7% |
| 17 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) Anthropic | May 2026 | 57.3 | 74.3 | — | $10.00 | $5.00 | $25.00 | 0 | 0.00 | 0.00 | 92.0% | 48.7% | 53.5% | 73.0% | 62.2% | 94.4% | 58.3% | — | — | — | — | 84.6% | — | 34.2% |
| 18 | Muse Spark 1.2 (xhigh) Meta | Aug 2026 | 56.8 | 72.2 | — | $2.00 | $1.25 | $4.25 | 0 | 0.00 | 0.00 | 90.4% | 45.5% | 56.4% | 83.3% | — | — | — | — | — | — | — | 80.1% | — | 34.8% |
| 19 | GPT-5.6 Terra (max) OpenAI | Jul 2026 | 56.6 | 76.7 | — | $4.50 | $2.00 | $12.00 | 110 | 119.68 | 119.68 | 92.5% | 42.9% | 53.9% | 79.7% | 71.2% | 86.3% | 57.6% | — | — | — | — | 88.0% | — | 40.2% |
| 20 | GPT-5.5 (xhigh) OpenAI | Apr 2026 | 56.3 | 74.9 | — | $11.25 | $5.00 | $30.00 | 0 | 0.00 | 0.00 | 93.5% | 45.8% | 56.1% | 79.0% | 75.9% | 93.9% | 60.6% | — | — | — | — | 84.3% | — | 39.0% |
| 21 | Gemini 3.7 Flash (high) Google | Aug 2026 | 56.0 | 76.1 | — | $1.50 | $0.75 | $3.75 | 345 | 7.57 | 7.57 | 94.5% | 47.9% | 56.8% | 80.0% | — | — | — | — | — | — | — | 85.8% | — | 32.8% |
| 22 | Qwen3.8-Flash-Next Alibaba | Aug 2026 | 55.8 | 73.1 | — | $0.23 | $0.15 | $0.47 | 83 | 1.22 | 25.36 | 92.3% | 38.0% | 46.9% | 77.0% | — | — | — | — | — | — | — | 86.1% | — | 45.4% |
| 23 | Grok 4.5 (high) SpaceXAI | Jul 2026 | 55.8 | 72.4 | — | $3.00 | $2.00 | $6.00 | 54 | 13.12 | 13.12 | 93.1% | 42.7% | 54.1% | 74.0% | — | — | — | — | — | — | — | 81.6% | — | 42.1% |
| 24 | GPT-5.6 Sol (medium) OpenAI | Jul 2026 | 55.6 | 76.3 | — | $8.00 | $4.00 | $20.00 | 78 | 5.35 | 5.35 | 92.6% | 42.2% | 56.5% | 74.3% | 69.6% | 81.0% | 62.9% | — | — | — | — | 86.1% | — | 36.5% |
| 25 | Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Anthropic | Jun 2026 | 55.3 | 71.5 | — | $4.00 | $2.00 | $10.00 | 91 | 121.77 | 121.77 | 91.1% | 41.3% | 53.6% | 77.0% | — | — | — | — | — | — | — | 80.5% | — | 37.3% |
| 26 | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) Anthropic | Apr 2026 | 55.0 | 73.6 | — | $10.00 | $5.00 | $25.00 | 0 | 0.00 | 0.00 | 91.4% | 42.3% | 54.5% | 75.3% | 58.6% | 88.6% | 51.5% | — | — | — | — | 83.1% | — | 34.6% |
| 27 | GPT-5.5 (high) OpenAI | Apr 2026 | 54.7 | 71.6 | — | $11.25 | $5.00 | $30.00 | 0 | 0.00 | 0.00 | 93.2% | 45.0% | 55.9% | 79.0% | 71.6% | 93.0% | 59.8% | — | — | — | — | 79.4% | — | 36.7% |
| 28 | Gemini 3.7 Flash (medium) Google | Aug 2026 | 53.4 | 71.5 | — | $1.50 | $0.75 | $3.75 | 327 | 4.42 | 4.42 | 92.1% | 39.0% | 57.9% | 81.0% | — | — | — | — | — | — | — | 78.3% | — | 35.5% |
| 29 | DeepSeek V4 Pro 0813 (Reasoning, Max Effort) DeepSeek | Aug 2026 | 53.2 | 68.8 | — | $1.98 | $1.32 | $3.96 | 62 | 1.15 | 33.25 | 92.8% | 41.0% | 49.2% | 75.3% | — | — | — | — | — | — | — | 78.7% | — | 39.6% |
| 30 | Muse Spark 1.1 (xhigh) Meta | Jul 2026 | 53.2 | 71.3 | — | $2.00 | $1.25 | $4.25 | 0 | 0.00 | 0.00 | 89.8% | 46.2% | 58.2% | 81.3% | — | — | — | — | — | — | — | 77.9% | — | 31.8% |
Independent benchmarks & pricing of LLMs — we publish them here, we don’t produce them. For further analysis and methodology, see artificialanalysis.ai. Updated 1 Sep 2026.
How to read it
Three things that trip people up when they first use this table.
Price is per million tokens
Input and output are priced separately, and output is usually the expensive half. A workflow that reads a lot and writes a little costs far less than the headline number suggests.
Two different kinds of speed
Time to first token is what a person waiting on a chat reply feels. Tokens per second is what matters for a batch job running overnight. Optimising for the wrong one wastes money.
Quality scores are a shortlist, not a verdict
Aggregate indices are a good way to narrow twenty models to three. They are a poor way to choose between those three — for that, test them on your actual task.
Building a workflow? We pick the model for you
We cost and benchmark every automation before it ships — so it runs on the cheapest model that actually holds up on your task.