Cortex Lab

LLM performance leaderboard

Price, output speed and quality across every major model, compared side by side. This is the table we open before deciding what a workflow should run on. The data is published by Artificial Analysis — we host it here for convenience and credit it in full.

120 models
Quality index Cost $/1M tokens Speed Benchmarks
# Model
1 Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic Jul 2026 63.1 78.0 $10.00 $5.00 $25.00 54 34.57 34.57 93.2% 54.9% 55.7% 75.7% 89.1% 42.1%
2 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Anthropic Jul 2026 62.5 77.0 $10.00 $5.00 $25.00 54 18.21 18.21 93.7% 54.4% 55.0% 76.3% 88.0% 43.3%
3 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Anthropic Jun 2026 62.1 76.5 $20.00 $10.00 $50.00 62 54.13 54.13 92.6% 55.5% 60.2% 76.7% 63.5% 98.5% 62.9% 84.6% 38.1%
4 Claude Opus 5 (Adaptive Reasoning, High Effort) Anthropic Jul 2026 61.5 76.5 $10.00 $5.00 $25.00 53 11.33 11.33 93.7% 52.8% 54.3% 76.3% 87.6% 44.7%
5 GPT-5.6 Sol (max) OpenAI Jul 2026 60.9 77.4 $8.00 $4.00 $20.00 84 50.56 50.56 94.1% 49.5% 56.1% 77.7% 72.7% 85.1% 65.9% 88.0% 44.3%
6 Grok 4.6 (high) SpaceXAI Aug 2026 60.9 76.8 $3.00 $2.00 $6.00 58 36.59 36.59 94.9% 42.9% 53.6% 75.0% 88.4% 50.7%
7 Grok 4.6 (xhigh) SpaceXAI Aug 2026 60.0 75.9 $3.00 $2.00 $6.00 60 38.68 38.68 93.5% 44.1% 51.6% 75.7% 88.0% 43.3%
8 Kimi K3 (max) Kimi Jul 2026 59.7 76.2 $6.00 $3.00 $15.00 36 3.56 58.53 93.5% 46.9% 58.7% 82.7% 85.0% 46.0%
9 GLM-5.3 (max) Z AI Aug 2026 59.5 74.8 $2.15 $1.40 $4.40 75 1.61 28.20 91.7% 42.3% 56.5% 76.3% 83.9% 50.3%
10 GPT-5.6 Sol (xhigh) OpenAI Jul 2026 59.0 78.3 $8.00 $4.00 $20.00 82 23.49 23.49 93.1% 47.3% 56.0% 76.3% 71.0% 84.8% 61.4% 89.5% 38.1%
11 Grok 4.6 (medium) SpaceXAI Aug 2026 59.0 74.4 $3.00 $2.00 $6.00 57 31.70 31.70 93.5% 42.1% 54.6% 72.7% 84.3% 44.3%
12 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Anthropic Jul 2026 58.6 74.3 $10.00 $5.00 $25.00 51 8.65 8.65 91.9% 51.3% 50.7% 78.7% 86.1% 38.6%
13 Qwen3.8 Max Alibaba Aug 2026 58.1 71.8 $3.00 $2.00 $6.00 30 1.76 68.67 92.7% 43.0% 52.9% 74.3% 81.3% 51.3%
14 Qwen3.8 2.4T A95B Alibaba Aug 2026 57.7 71.9 $3.00 $2.00 $6.00 31 1.91 65.63 93.5% 42.4% 51.6% 75.3% 82.0% 49.1%
15 GLM-5.3-Flash Z AI Aug 2026 57.5 71.5 $0.24 $0.15 $0.50 50 1.11 41.02 91.2% 39.9% 46.1% 78.0% 84.3% 47.2%
16 GPT-5.6 Sol (high) OpenAI Jul 2026 57.3 77.2 $8.00 $4.00 $20.00 78 14.11 14.11 92.8% 46.0% 56.9% 75.3% 69.2% 83.3% 62.1% 87.3% 36.7%
17 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) Anthropic May 2026 57.3 74.3 $10.00 $5.00 $25.00 0 0.00 0.00 92.0% 48.7% 53.5% 73.0% 62.2% 94.4% 58.3% 84.6% 34.2%
18 Muse Spark 1.2 (xhigh) Meta Aug 2026 56.8 72.2 $2.00 $1.25 $4.25 0 0.00 0.00 90.4% 45.5% 56.4% 83.3% 80.1% 34.8%
19 GPT-5.6 Terra (max) OpenAI Jul 2026 56.6 76.7 $4.50 $2.00 $12.00 110 119.68 119.68 92.5% 42.9% 53.9% 79.7% 71.2% 86.3% 57.6% 88.0% 40.2%
20 GPT-5.5 (xhigh) OpenAI Apr 2026 56.3 74.9 $11.25 $5.00 $30.00 0 0.00 0.00 93.5% 45.8% 56.1% 79.0% 75.9% 93.9% 60.6% 84.3% 39.0%
21 Gemini 3.7 Flash (high) Google Aug 2026 56.0 76.1 $1.50 $0.75 $3.75 345 7.57 7.57 94.5% 47.9% 56.8% 80.0% 85.8% 32.8%
22 Qwen3.8-Flash-Next Alibaba Aug 2026 55.8 73.1 $0.23 $0.15 $0.47 83 1.22 25.36 92.3% 38.0% 46.9% 77.0% 86.1% 45.4%
23 Grok 4.5 (high) SpaceXAI Jul 2026 55.8 72.4 $3.00 $2.00 $6.00 54 13.12 13.12 93.1% 42.7% 54.1% 74.0% 81.6% 42.1%
24 GPT-5.6 Sol (medium) OpenAI Jul 2026 55.6 76.3 $8.00 $4.00 $20.00 78 5.35 5.35 92.6% 42.2% 56.5% 74.3% 69.6% 81.0% 62.9% 86.1% 36.5%
25 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Anthropic Jun 2026 55.3 71.5 $4.00 $2.00 $10.00 91 121.77 121.77 91.1% 41.3% 53.6% 77.0% 80.5% 37.3%
26 Claude Opus 4.7 (Adaptive Reasoning, Max Effort) Anthropic Apr 2026 55.0 73.6 $10.00 $5.00 $25.00 0 0.00 0.00 91.4% 42.3% 54.5% 75.3% 58.6% 88.6% 51.5% 83.1% 34.6%
27 GPT-5.5 (high) OpenAI Apr 2026 54.7 71.6 $11.25 $5.00 $30.00 0 0.00 0.00 93.2% 45.0% 55.9% 79.0% 71.6% 93.0% 59.8% 79.4% 36.7%
28 Gemini 3.7 Flash (medium) Google Aug 2026 53.4 71.5 $1.50 $0.75 $3.75 327 4.42 4.42 92.1% 39.0% 57.9% 81.0% 78.3% 35.5%
29 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) DeepSeek Aug 2026 53.2 68.8 $1.98 $1.32 $3.96 62 1.15 33.25 92.8% 41.0% 49.2% 75.3% 78.7% 39.6%
30 Muse Spark 1.1 (xhigh) Meta Jul 2026 53.2 71.3 $2.00 $1.25 $4.25 0 0.00 0.00 89.8% 46.2% 58.2% 81.3% 77.9% 31.8%

Independent benchmarks & pricing of LLMs — we publish them here, we don’t produce them. For further analysis and methodology, see artificialanalysis.ai. Updated 1 Sep 2026.

How to read it

Three things that trip people up when they first use this table.

Price is per million tokens

Input and output are priced separately, and output is usually the expensive half. A workflow that reads a lot and writes a little costs far less than the headline number suggests.

Two different kinds of speed

Time to first token is what a person waiting on a chat reply feels. Tokens per second is what matters for a batch job running overnight. Optimising for the wrong one wastes money.

Quality scores are a shortlist, not a verdict

Aggregate indices are a good way to narrow twenty models to three. They are a poor way to choose between those three — for that, test them on your actual task.

Building a workflow? We pick the model for you

We cost and benchmark every automation before it ships — so it runs on the cheapest model that actually holds up on your task.