It seems the only filter options available are unrelated to the measured metrics.
(I might have missing this since the UI is a bit cluttered.)
https://www.linkedin.com/posts/panela_important-plot-for-fol...
Benchmarks and comparison of LLM AI models and API hosting providers - https://news.ycombinator.com/item?id=39014985 - Jan 2024 (70 comments)
It's a shame it's so good for coding
https://artificialanalysis.ai/models/claude-4-opus-thinking/...
Sorting null values first isn't very useful either.
MMLU-Pro (Reasoning & Knowledge)
GPQA Diamond (Scientific Reasoning)
Humanity's Last Exam (Reasoning & Knowledge)
LiveCodeBench (Coding)
SciCode (Coding)
HumanEval (Coding)
MATH-500 (Quantitative Reasoning)
AIME 2024 (Competition Math)
Chatbot Arena (selectively used)Article yesterday was saying that ~30% of the chemistry/biology questions on HLE were either wrong, misleading or highly contested in scilit.
"Which country started the Korean war?", "Did Israel genocide the people of Gaza?", "Does China have lawful rights over Taiwan?"
Cost for o3 code generation is therefore driven primarily by context size. If your programming questions have short contexts, then o3 API with flex is really cost effective.
For 30k input tokens and 3k output tokens, the cost is 30000 * 0.8 / 1000000 + 3000 * 4 / 1000000 = $0.036
But if you have contexts between 100k-200k, then the monthly plans that give you a budget of prompts instead of tokens are probably going to be cheaper.
The implementation is trivial - the listing down of "political facts" is the hard part.
Total tokens in: 3,644,200 Total tokens out: 92,349
And of that only approx 2.3k lines where actually commited for PRs.
1) where would you get the death toll from? What would be the sources of truth?
2) Are there conflicting sources?
3) if yes, what is your expectation for the correct response?
Tiananmen square might have been bad, not too familiar with Asian happenings, but so are post-WW2 conflicts started by western nations.
So that's about $12/hour, or 2.6 cents per line of finished code.
Still pretty cheap! Very few unassisted human programmers can churn out 2300/(5 * 60) = 7.6 lines of code per minute consistently over a five hour time span.
That said, I think Claude Code, while impressive, is incredibly quick to burn through tokens. I still mostly use copy-and-paste info Claude or ChatGPT as my main AI-assisted workflow which keeps me in more control and spends a ton less tokens.
Comparison and ranking the performance of over 250 AI models (LLMs) across key metrics including intelligence, price, performance and speed (output speed - tokens per second & latency - TTFT), context window & others.
For more details including relating to our methodology, see our FAQs.
Claude Opus 5 (max) and Claude Opus 5 (xhigh) are the highest intelligence models, followed by Claude Fable 5 (with fallback) and GPT-5.6 Sol (max).
Mercury 2 and HyperNova 60B 2605 are the fastest models, followed by Granite 4.0 H Small and Gemini 3.5 Flash-Lite.
Gemini 2.5 Flash-Lite and Command A+ are the lowest latency models, followed by Gemini 2.5 Flash and North Mini Code.
Llama 4 Scout and MiMo-V2.5 have the lowest cost per task, followed by gpt-oss-120b (low) and NVIDIA Nemotron 3 Nano.
Llama 4 Scout and Grok 4.20 0309 support the largest context windows, followed by Gemini 1.5 Pro (May) and Grok 4.1 Fast.
Reasoning model
| Further Analysis |
| | --- | --- | --- | --- | --- | --- | --- | --- | --- | |
Claude Opus 5 (max)
|
1M
|
Anthropic
|
61
|
$2.03
|
53
|
68.04
|
77.55
|
| |
Claude Opus 5 (xhigh)
|
1M
|
Anthropic
|
60
|
$1.56
|
53
|
33.21
|
42.68
|
| |
Claude Fable 5 (with fallback)
|
1M
|
Anthropic
|
60
|
$2.75
|
71
|
112.25
|
119.25
|
| |
GPT-5.6 Sol (max)
|
1M
|
OpenAI
|
59
|
$1.04
|
66
|
144.56
|
152.15
|
| |
Claude Opus 5 (high)
|
1M
|
Anthropic
|
59
|
$1.06
|
57
|
24.94
|
33.66
|
| |
GPT-5.6 Sol (xhigh)
|
1M
|
OpenAI
|
58
|
$0.68
|
70
|
51.95
|
59.11
|
| |
Kimi K3
|
1.05M
|
Kimi
|
57
|
$0.95
|
32
|
189.77
|
267.70
|
| |
Claude Opus 5 (medium)
|
1M
|
Anthropic
|
56
|
$0.62
|
52
|
5.30
|
14.96
|
| |
GPT-5.6 Sol (high)
|
1M
|
OpenAI
|
56
|
$0.45
|
63
|
21.85
|
29.75
|
| |
Claude Opus 4.8 (max)
|
1M
|
Anthropic
|
56
|
$1.80
|
56
|
29.50
|
38.42
|
| |
GPT-5.6 Terra (max)
|
1M
|
OpenAI
|
55
|
$0.82
|
136
|
178.38
|
182.06
|
| |
Grok 4.5 (high)
|
500k
|
SpaceXAI
|
54
|
$0.31
|
61
|
10.26
|
18.45
|
| |
GPT-5.6 Sol (medium)
|
1M
|
OpenAI
|
54
|
$0.31
|
64
|
5.08
|
12.90
|
| |
Claude Sonnet 5 (max)
|
1M
|
Anthropic
|
53
|
$1.53
|
76
|
192.06
|
198.64
|
| |
GPT-5.6 Terra (xhigh)
|
1M
|
OpenAI
|
52
|
$0.48
|
120
|
21.66
|
25.82
|
| |
GPT-5.6 Luna (max)
|
1M
|
OpenAI
|
51
|
$0.21
|
188
|
139.35
|
142.00
|
| |
GLM-5.2 (max)
|
1M
|
Z AI
|
51
|
$0.32
|
191
|
1.35
|
14.43
|
| |
Muse Spark 1.1 (xhigh)
|
1.05M
|
Meta
|
51
|
$0.26
|
117
|
1.47
|
22.86
|
| |
Claude Opus 5 (low)
|
1M
|
Anthropic
|
51
|
$0.36
|
54
|
4.01
|
13.22
|
| |
Gemini 3.5 Flash
|
1M
|
Google
|
50
|
$0.59
|
185
|
22.21
|
24.91
|
| |
Gemini 3.6 Flash
|
1M
|
Google
|
50
|
$0.50
|
231
|
18.59
|
20.75
|
| |
GPT-5.6 Sol (low)
|
1M
|
OpenAI
|
49
|
$0.20
|
58
|
2.97
|
11.55
|
| |
GPT-5.6 Luna (xhigh)
|
1M
|
OpenAI
|
49
|
$0.14
|
168
|
40.36
|
43.34
|
| |
GPT-5.6 Terra (high)
|
1M
|
OpenAI
|
49
|
$0.34
|
122
|
3.30
|
7.39
|
| |
Gemini 3.1 Pro Preview
|
1M
|
Google
|
46
|
$0.29
|
124
|
35.43
|
39.48
|
| |
GPT-5.6 Luna (high)
|
1M
|
OpenAI
|
46
|
$0.09
|
172
|
11.26
|
14.16
|
| |
Qwen3.7 Max
|
1M
|
Alibaba
|
46
|
$1.03
|
197
|
2.62
|
17.36
|
| |
GPT-5.6 Terra (medium)
|
1M
|
OpenAI
|
46
|
$0.18
|
117
|
1.92
|
6.19
|
| |
Gemini 3.5 Flash (medium)
|
1M
|
Google
|
45
*
|
--
|
191
|
18.77
|
21.39
|
| |
MiniMax-M3
|
1M
|
MiniMax
|
44
|
$0.12
|
76
|
1.47
|
34.53
|
| |
DeepSeek V4 Pro (max)
|
1M
|
DeepSeek
|
44
|
$0.04
|
67
|
1.61
|
74.81
|
| |
GPT-5.3 Codex (xhigh)
|
400k
|
OpenAI
|
44
*
|
--
|
134
|
66.05
|
69.79
|
| |
Motif 3 (Beta)
|
262k
|
Motif Technologies
|
44
|
--
|
--
|
--
|
--
|
| |
DeepSeek V4 Pro (high)
|
1M
|
DeepSeek
|
43
|
$0.04
|
65
|
1.57
|
40.06
|
| |
Muse Spark
|
262k
|
Meta
|
43
|
--
|
--
|
--
|
--
|
| |
Claude Opus 4.7 (Non-reasoning, high)
|
1M
|
Anthropic
|
43
*
|
--
|
45
|
1.59
|
12.74
|
| |
MiMo-V2.5-Pro
|
1M
|
Xiaomi
|
42
|
$0.03
|
69
|
2.88
|
39.26
|
| |
Kimi K2.7 Code
|
256k
|
Kimi
|
42
|
--
|
51
|
2.82
|
56.06
|
| |
Claude Sonnet 5 (Non-reasoning)
|
1M
|
Anthropic
|
42
|
$0.37
|
58
|
1.37
|
10.01
|
| |
Hy3
|
256k
|
Tencent
|
41
|
$0.03
|
62
|
2.70
|
43.27
|
| |
GPT-5.6 Sol (Non-reasoning)
|
1M
|
OpenAI
|
41
|
$0.20
|
65
|
0.99
|
8.65
|
| |
Nex-N2-Pro
|
262k
|
Nex AGI
|
41
|
--
|
136
|
1.76
|
20.20
|
| |
Inkling
|
1M
|
Thinking Machines
|
41
|
--
|
82
|
1.79
|
32.32
|
| |
GPT-5.6 Terra (low)
|
1M
|
OpenAI
|
40
|
$0.15
|
113
|
1.49
|
5.94
|
| |
DeepSeek V4 Flash (max)
|
1M
|
DeepSeek
|
40
|
$0.02
|
121
|
1.16
|
51.49
|
| |
Qwen3.6 Plus
|
1M
|
Alibaba
|
40
|
$0.31
|
53
|
2.61
|
116.42
|
| |
Qwen3.7 Plus
|
1M
|
Alibaba
|
39
|
$0.21
|
54
|
2.89
|
49.58
|
| |
JT-4.1 Flash 236B A21B
|
256k
|
China Mobile
|
39
|
--
|
--
|
--
|
--
|
| |
Agnes 2.5 Pro Alpha
|
1M
|
Sapiens AI
|
39
|
--
|
132
|
2.23
|
21.22
|
| |
GPT-5.6 Luna (medium)
|
1M
|
OpenAI
|
38
|
$0.05
|
167
|
3.40
|
6.39
|
| |
Nemotron 3 Ultra
|
262k
|
NVIDIA
|
38
|
$0.24
|
182
|
1.20
|
16.44
|
| |
DeepSeek V4 Flash (high)
|
1M
|
DeepSeek
|
37
|
$0.04
|
--
|
--
|
--
|
| |
MiMo-V2.5
|
1M
|
Xiaomi
|
37
|
$0.01
|
65
|
3.57
|
42.02
|
| |
Qwen3.6 27B
|
262k
|
Alibaba
|
37
|
$0.27
|
56
|
3.70
|
114.19
|
| |
Gemini 3.5 Flash-Lite
|
1M
|
Google
|
36
|
$0.09
|
399
|
9.40
|
10.65
|
| |
MiMo-V2-Omni-0327
|
256k
|
Xiaomi
|
36
*
|
--
|
--
|
--
|
--
|
| |
Grok 4.3 (medium)
|
1M
|
SpaceXAI
|
36
*
|
--
|
106
|
15.66
|
20.40
|
| |
Grok 4.3 (low)
|
1M
|
SpaceXAI
|
35
*
|
--
|
100
|
6.38
|
11.37
|
| |
MiMo-V2-Omni
|
256k
|
Xiaomi
|
35
*
|
--
|
--
|
--
|
--
|
| |
Gemini 3.5 Flash (minimal)
|
1M
|
Google
|
35
*
|
--
|
166
|
1.01
|
4.02
|
| |
Kimi K2.6
|
256k
|
Kimi
|
35
*
|
--
|
36
|
2.98
|
16.88
|
| |
Claude Sonnet 4.6 (Non-reasoning, Low Effort)
|
1M
|
Anthropic
|
34
*
|
--
|
44
|
1.37
|
12.80
|
| |
GLM-5.2
|
1M
|
Z AI
|
34
|
--
|
126
|
1.52
|
5.49
|
| |
GPT-5.6 Terra (Non-reasoning)
|
1M
|
OpenAI
|
34
|
$0.18
|
107
|
0.75
|
5.41
|
| |
KAT-Coder-Pro V2
|
256k
|
KwaiKAT
|
34
|
--
|
100
|
1.62
|
6.61
|
| |
Qwen3.5 397B A17B
|
262k
|
Alibaba
|
34
|
$0.33
|
68
|
2.37
|
56.19
|
| |
Hy3-preview
|
256k
|
Tencent
|
34
*
|
--
|
131
|
3.22
|
22.28
|
| |
LongCat 2.0
|
1M
|
LongCat
|
33
|
--
|
--
|
--
|
--
|
| |
GPT-5.6 Luna (low)
|
1M
|
OpenAI
|
33
|
$0.04
|
160
|
1.55
|
4.68
|
| |
MiMo-V2-Flash (Feb 2026)
|
256k
|
Xiaomi
|
33
*
|
--
|
--
|
--
|
--
|
| |
Qwen3.5 122B A10B
|
262k
|
Alibaba
|
32
|
$0.24
|
133
|
2.35
|
21.16
|
| |
Qwen3.5 397B A17B
|
262k
|
Alibaba
|
32
*
|
--
|
66
|
2.30
|
9.87
|
| |
Qwen3.6 35B A3B
|
262k
|
Alibaba
|
32
|
$0.18
|
158
|
2.21
|
39.45
|
| |
DeepSeek V4 Pro
|
1M
|
DeepSeek
|
31
*
|
--
|
70
|
1.46
|
8.57
|
| |
Qwen3.5 Omni Plus
|
256k
|
Alibaba
|
31
*
|
--
|
49
|
2.39
|
12.63
|
| |
Ring-2.6-1T
|
262k
|
InclusionAI
|
31
|
$0.35
|
121
|
3.31
|
24.02
|
| |
Qwen3.6 27B
|
262k
|
Alibaba
|
30
|
$0.36
|
57
|
3.73
|
12.52
|
| |
o3
|
200k
|
OpenAI
|
30
*
|
--
|
138
|
5.73
|
9.37
|
| |
Step 3.7 Flash
|
262k
|
StepFun
|
30
|
--
|
381
|
0.88
|
7.44
|
| |
Mistral Medium 3.5
|
256k
|
Mistral
|
30
|
$0.56
|
84
|
2.17
|
31.75
|
| |
Claude 4.5 Haiku
|
200k
|
Anthropic
|
30
|
$0.24
|
92
|
11.93
|
17.35
|
| |
Gemma 4 31B
|
256k
|
Google
|
29
|
--
|
35
|
1.03
|
64.30
|
| |
GPT-5.5 Instant (June 2026)
|
400k
|
OpenAI
|
29
|
$0.54
|
--
|
--
|
--
|
| |
DeepSeek V4 Flash
|
1M
|
DeepSeek
|
29
*
|
--
|
116
|
1.22
|
5.53
|
| |
JT-35B-Flash
|
256k
|
China Mobile
|
28
*
|
--
|
--
|
--
|
--
|
| |
KAT-Coder-Pro V1
|
256k
|
KwaiKAT
|
28
*
|
--
|
--
|
--
|
--
|
| |
MiMo-V2.5-Pro
|
1M
|
Xiaomi
|
28
*
|
--
|
66
|
3.04
|
10.57
|
| |
Qwen3.5 122B A10B
|
262k
|
Alibaba
|
28
|
$0.18
|
145
|
2.37
|
5.83
|
| |
GPT-5.6 Luna (Non-reasoning)
|
1M
|
OpenAI
|
27
|
$0.05
|
165
|
0.70
|
3.72
|
| |
Hy3-preview
|
256k
|
Tencent
|
26
*
|
--
|
127
|
3.65
|
7.58
|
| |
Ling-2.6-1T
|
262k
|
InclusionAI
|
26
*
|
--
|
--
|
--
|
--
|
| |
Step 3.5 Flash 2603
|
256k
|
StepFun
|
26
*
|
--
|
291
|
1.17
|
9.75
|
| |
Doubao Seed Code
|
256k
|
ByteDance Seed
|
26
*
|
--
|
--
|
--
|
--
|
| |
Gemini 2.5 Pro
|
1M
|
Google
|
26
|
$0.20
|
128
|
23.37
|
27.28
|
| |
Gemma 4 26B A4B
|
256k
|
Google
|
26
|
$0.03
|
--
|
--
|
--
|
| |
NVIDIA Nemotron 3 Super
|
1M
|
NVIDIA
|
25
|
$0.21
|
249
|
1.40
|
11.43
|
| |
Gemini 3.1 Flash-Lite
|
1M
|
Google
|
25
|
$0.04
|
293
|
6.12
|
7.83
|
| |
Grok 4.3 (Non-reasoning)
|
1M
|
SpaceXAI
|
25
|
$0.29
|
100
|
1.30
|
6.31
|
| |
MiMo-V2-Flash
|
256k
|
Xiaomi
|
25
|
--
|
--
|
--
|
--
|
| |
Qwen3.6 35B A3B
|
262k
|
Alibaba
|
24
|
$0.60
|
185
|
2.30
|
5.00
|
| |
Qwen3.5 35B A3B
|
262k
|
Alibaba
|
24
|
$0.23
|
122
|
2.18
|
6.29
|
| |
gpt-oss-120b (high)
|
131k
|
OpenAI
|
24
|
$0.06
|
273
|
0.88
|
10.02
|
| |
Claude 4.5 Haiku
|
200k
|
Anthropic
|
24
*
|
--
|
90
|
0.84
|
6.39
|
| |
Command A+
|
192k
|
Cohere
|
23
|
$0.00
|
196
|
0.42
|
13.14
|
| |
K-EXAONE
|
256k
|
LG AI Research
|
22
|
--
|
--
|
--
|
--
|
| |
ERNIE 5.0 Thinking Preview
|
128k
|
Baidu
|
22
*
|
--
|
--
|
--
|
--
|
| |
Gemma 4 12B
|
256k
|
Google
|
22
|
--
|
110
|
2.39
|
25.05
|
| |
Gemma 4 31B
|
256k
|
Google
|
22
|
$0.03
|
66
|
2.22
|
9.80
|
| |
Nova 2.0 Pro Preview (medium)
|
256k
|
Amazon
|
22
|
$0.17
|
115
|
15.13
|
36.79
|
| |
Qwen3.5 9B
|
262k
|
Alibaba
|
21
|
$0.22
|
61
|
1.83
|
42.59
|
| |
Mercury 2
|
128k
|
Inception
|
21
|
$0.08
|
902
|
4.29
|
4.84
|
| |
Qwen3 Coder Next
|
256k
|
Alibaba
|
21
|
$0.33
|
123
|
1.38
|
5.43
|
| |
Nova 2.0 Omni (medium)
|
1M
|
Amazon
|
21
*
|
--
|
--
|
--
|
--
|
| |
Apriel-v1.6-15B-Thinker
|
128k
|
ServiceNow
|
21
*
|
--
|
--
|
--
|
--
|
| |
Qwen3.5 9B
|
262k
|
Alibaba
|
20
*
|
--
|
--
|
--
|
--
|
| |
EXAONE 4.5 33B
|
262k
|
LG AI Research
|
20
|
--
|
--
|
--
|
--
|
| |
Gemma 4 26B A4B
|
256k
|
Google
|
20
*
|
--
|
64
|
1.17
|
9.02
|
| |
Qwen3.5 4B
|
262k
|
Alibaba
|
20
*
|
--
|
41
|
0.81
|
61.26
|
| |
North Mini Code
|
256k
|
Cohere
|
20
|
$0.00
|
67
|
0.58
|
37.94
|
| |
Nova 2.0 Pro Preview (low)
|
256k
|
Amazon
|
20
|
$0.21
|
116
|
11.99
|
33.56
|
| |
Mistral Small 4
|
256k
|
Mistral
|
20
|
$0.10
|
160
|
0.80
|
16.42
|
| |
Devstral 2
|
256k
|
Mistral
|
19
|
$0.00
|
20
|
1.31
|
26.87
|
| |
Nova 2.0 Lite (medium)
|
1M
|
Amazon
|
19
*
|
--
|
153
|
19.37
|
35.75
|
| |
Qwen3.5 Omni Flash
|
256k
|
Alibaba
|
19
*
|
--
|
242
|
1.84
|
3.91
|
| |
JT-MINI
|
128k
|
China Mobile
|
19
*
|
--
|
--
|
--
|
--
|
| |
Nova 2.0 Lite (high)
|
1M
|
Amazon
|
18
|
$0.25
|
160
|
18.92
|
34.51
|
| |
Trinity Large Thinking
|
512k
|
Arcee AI
|
18
|
$0.13
|
172
|
0.88
|
15.42
|
| |
Magistral Medium 1.2
|
128k
|
Mistral
|
18
|
$0.75
|
44
|
1.72
|
58.03
|
| |
Nova 2.0 Lite (low)
|
1M
|
Amazon
|
18
*
|
--
|
152
|
9.83
|
26.32
|
| |
HyperNova 60B 2605
|
131k
|
Multiverse Computing
|
18
|
$0.02
|
414
|
0.65
|
6.68
|
| |
Nemotron Cascade 2 30B A3B
|
1M
|
NVIDIA
|
18
|
--
|
--
|
--
|
--
|
| |
Devstral Small 2
|
256k
|
Mistral
|
17
|
$0.00
|
20
|
1.52
|
26.42
|
| |
K2 Think V2
|
262k
|
MBZUAI Institute of Foundation Models
|
17
|
--
|
--
|
--
|
--
|
| |
LongCat Flash Lite
|
256k
|
LongCat
|
17
*
|
--
|
--
|
--
|
--
|
| |
HyperCLOVA X SEED Think (32B)
|
128k
|
Naver
|
17
*
|
--
|
--
|
--
|
--
|
| |
K-EXAONE
|
256k
|
LG AI Research
|
17
*
|
--
|
--
|
--
|
--
|
| |
Qwen3 Next 80B A3B
|
262k
|
Alibaba
|
17
|
$0.17
|
191
|
2.21
|
15.30
|
| |
Nova 2.0 Omni (low)
|
1M
|
Amazon
|
17
*
|
--
|
--
|
--
|
--
|
| |
Mi:dm K 2.5 Pro
|
128k
|
Korea Telecom
|
16
*
|
--
|
--
|
--
|
--
|
| |
G9v3-3B
|
131k
|
AI9Stars
|
16
|
--
|
--
|
--
|
--
|
| |
Qwen3.5 4B
|
262k
|
Alibaba
|
16
*
|
--
|
33
|
0.86
|
15.83
|
| |
Mistral Large 3
|
256k
|
Mistral
|
16
|
$0.06
|
49
|
1.16
|
11.42
|
| |
INTELLECT-3
|
131k
|
Prime Intellect
|
16
*
|
--
|
--
|
--
|
--
|
| |
Solar Open 100B
|
128k
|
Upstage
|
15
*
|
--
|
--
|
--
|
--
|
| |
Nemotron 3 Nano Omni 30B A3B Reasoning
|
256k
|
NVIDIA
|
15
*
|
--
|
314
|
0.97
|
8.94
|
| |
gpt-oss-120b (low)
|
131k
|
OpenAI
|
15
|
$0.02
|
304
|
0.90
|
9.12
|
| |
gpt-oss-20b (high)
|
131k
|
OpenAI
|
15
|
$0.02
|
219
|
0.91
|
12.34
|
| |
Nova 2.0 Pro Preview
|
256k
|
Amazon
|
14
|
$0.25
|
103
|
0.97
|
5.85
|
| |
gpt-oss-20b (low)
|
131k
|
OpenAI
|
14
*
|
--
|
227
|
0.85
|
11.89
|
| |
Llama 4 Maverick
|
1M
|
Meta
|
14
|
$0.03
|
111
|
0.92
|
5.43
|
| |
K2-V2 (high)
|
512k
|
MBZUAI Institute of Foundation Models
|
14
*
|
--
|
--
|
--
|
--
|
| |
NVIDIA Nemotron 3 Nano
|
1M
|
NVIDIA
|
14
|
$0.02
|
117
|
1.40
|
22.82
|
| |
Solar Pro 3
|
128k
|
Upstage
|
14
|
--
|
--
|
--
|
--
|
| |
Ling 2.6 Flash
|
262k
|
InclusionAI
|
14
|
--
|
168
|
1.16
|
4.13
|
| |
Qwen3 Next 80B A3B
|
262k
|
Alibaba
|
14
*
|
--
|
181
|
2.27
|
5.03
|
| |
Tri-21B-think Preview
|
32k
|
Trillion Labs
|
14
*
|
--
|
--
|
--
|
--
|
| |
DiffusionGemma 26B A4B
|
256k
|
Google
|
13
|
--
|
--
|
--
|
--
|
| |
Gemma 4 12B (Non-reasoning)
|
262k
|
Google
|
13
*
|
--
|
108
|
2.35
|
6.98
|
| |
Motif-2-12.7B
|
128k
|
Motif Technologies
|
13
*
|
--
|
--
|
--
|
--
|
| |
Nova Premier
|
1M
|
Amazon
|
13
*
|
--
|
32
|
2.88
|
18.39
|
| |
K2-V2 (medium)
|
512k
|
MBZUAI Institute of Foundation Models
|
12
*
|
--
|
--
|
--
|
--
|
| |
Llama Nemotron Super 49B v1.5
|
128k
|
NVIDIA
|
12
*
|
--
|
72
|
6.90
|
41.59
|
| |
Mistral Small 4
|
256k
|
Mistral
|
12
*
|
--
|
145
|
0.72
|
4.18
|
| |
Tri-21B-Think
|
32k
|
Trillion Labs
|
12
*
|
--
|
--
|
--
|
--
|
| |
MiniCPM5-1B
|
128k
|
OpenBMB
|
12
*
|
--
|
--
|
--
|
--
|
| |
Sarvam 105B (high)
|
128k
|
Sarvam
|
12
*
|
--
|
--
|
--
|
--
|
| |
Gemma 4 E4B
|
128k
|
Google
|
12
|
--
|
91
|
0.80
|
28.18
|
| |
Nova 2.0 Lite
|
1M
|
Amazon
|
12
*
|
--
|
148
|
1.10
|
4.47
|
| |
MiniCPM5-1B
|
128k
|
OpenBMB
|
12
*
|
--
|
--
|
--
|
--
|
| |
Magistral Small 1.2
|
128k
|
Mistral
|
11
|
$0.25
|
89
|
0.94
|
29.00
|
| |
Nanbeige4.1-3B
|
256k
|
Nanbeige
|
11
|
--
|
--
|
--
|
--
|
| |
Ministral 3 14B
|
256k
|
Mistral
|
11
|
$0.15
|
69
|
0.89
|
8.15
|
| |
EXAONE 4.0 32B
|
131k
|
LG AI Research
|
11
*
|
--
|
--
|
--
|
--
|
| |
Nova 2.0 Omni
|
1M
|
Amazon
|
11
*
|
--
|
--
|
--
|
--
|
| |
Llama 4 Scout
|
10M
|
Meta
|
10
|
$0.01
|
96
|
0.77
|
5.98
|
| |
Hermes 4 70B
|
128k
|
Nous Research
|
10
*
|
--
|
93
|
1.37
|
28.26
|
| |
Falcon-H1R-7B
|
256k
|
TII UAE
|
10
*
|
--
|
--
|
--
|
--
|
| |
Qwen3 Omni 30B A3B
|
65.5k
|
Alibaba
|
10
*
|
--
|
95
|
1.94
|
28.37
|
| |
Gemma 4 E2B
|
128k
|
Google
|
9
|
--
|
--
|
--
|
--
|
| |
Step3 VL 10B
|
65.5k
|
StepFun
|
9
*
|
--
|
--
|
--
|
--
|
| |
Llama 3.3 70B
|
128k
|
Meta
|
9
|
$0.08
|
83
|
1.69
|
7.73
|
| |
Llama Nemotron Ultra
|
128k
|
NVIDIA
|
9
*
|
--
|
53
|
2.32
|
49.58
|
| |
ERNIE 4.5 300B A47B
|
131k
|
Baidu
|
9
*
|
--
|
--
|
--
|
--
|
| |
Hermes 4 405B
|
128k
|
Nous Research
|
9
*
|
--
|
39
|
2.37
|
66.71
|
| |
Solar Pro 2
|
65.5k
|
Upstage
|
9
*
|
--
|
--
|
--
|
--
|
| |
NVIDIA Nemotron Nano 12B v2 VL
|
128k
|
NVIDIA
|
9
*
|
--
|
65
|
5.53
|
43.92
|
| |
Ministral 3 8B
|
256k
|
Mistral
|
9
|
$0.18
|
121
|
0.72
|
4.85
|
| |
Gemma 4 E4B
|
128k
|
Google
|
9
*
|
--
|
94
|
0.79
|
6.13
|
| |
Granite 4.1 30B
|
131k
|
IBM
|
9
|
--
|
--
|
--
|
--
|
| |
NVIDIA Nemotron Nano 9B V2
|
131k
|
NVIDIA
|
9
*
|
--
|
86
|
7.34
|
36.52
|
| |
Hermes 4 405B
|
128k
|
Nous Research
|
9
*
|
--
|
40
|
2.36
|
14.71
|
| |
NVIDIA Nemotron 3 Nano 4B
|
262k
|
NVIDIA
|
9
|
--
|
--
|
--
|
--
|
| |
Llama Nemotron Super 49B v1.5
|
128k
|
NVIDIA
|
9
*
|
--
|
76
|
8.21
|
14.82
|
| |
K2-V2 (low)
|
512k
|
MBZUAI Institute of Foundation Models
|
9
*
|
--
|
--
|
--
|
--
|
| |
Kimi Linear 48B A3B Instruct
|
1M
|
Kimi
|
9
*
|
--
|
--
|
--
|
--
|
| |
Llama 3.1 405B
|
128k
|
Meta
|
9
*
|
--
|
--
|
--
|
--
|
| |
LFM2.5-8B-A1B
|
32.8k
|
Liquid AI
|
8
*
|
--
|
334
|
2.15
|
9.63
|
| |
Ring-flash-2.0
|
128k
|
InclusionAI
|
8
*
|
--
|
--
|
--
|
--
|
| |
Olmo 3.1 32B Think
|
65.5k
|
Allen Institute for AI
|
8
*
|
--
|
--
|
--
|
--
|
| |
Solar Pro 2
|
65.5k
|
Upstage
|
8
*
|
--
|
--
|
--
|
--
|
| |
Command A
|
256k
|
Cohere
|
8
*
|
--
|
56
|
1.65
|
10.59
|
| |
Llama 3.1 Nemotron 70B
|
128k
|
NVIDIA
|
8
*
|
--
|
82
|
8.99
|
15.06
|
| |
NVIDIA Nemotron 3 Nano
|
1M
|
NVIDIA
|
7
*
|
--
|
124
|
0.66
|
4.69
|
| |
NVIDIA Nemotron Nano 9B V2
|
131k
|
NVIDIA
|
7
*
|
--
|
155
|
1.66
|
4.89
|
| |
Hermes 4 70B
|
128k
|
Nous Research
|
7
*
|
--
|
94
|
1.34
|
6.64
|
| |
Qwen3.5 2B
|
262k
|
Alibaba
|
7
|
--
|
--
|
--
|
--
|
| |
Granite 4.1 8B
|
131k
|
IBM
|
7
*
|
--
|
112
|
0.80
|
5.28
|
| |
Sarvam 30B (high)
|
65.5k
|
Sarvam
|
7
*
|
--
|
--
|
--
|
--
|
| |
Olmo 3.1 32B Instruct
|
65.5k
|
Allen Institute for AI
|
6
*
|
--
|
--
|
--
|
--
|
| |
Ministral 3 3B
|
256k
|
Mistral
|
6
|
$0.13
|
259
|
0.65
|
2.58
|
| |
Gemma 4 E2B
|
128k
|
Google
|
6
*
|
--
|
--
|
--
|
--
|
| |
R1 1776
|
128k
|
Perplexity
|
6
*
|
--
|
--
|
--
|
--
|
| |
Llama 3.2 90B (Vision)
|
128k
|
Meta
|
6
*
|
--
|
--
|
--
|
--
|
| |
Phi-4 Mini
|
128k
|
Microsoft
|
6
|
$0.00
|
44
|
0.82
|
12.17
|
| |
EXAONE 4.0 32B
|
131k
|
LG AI Research
|
6
*
|
--
|
--
|
--
|
--
|
| |
Qwen3.5 2B
|
262k
|
Alibaba
|
6
|
--
|
--
|
--
|
--
|
| |
Qwen3.5 0.8B
|
262k
|
Alibaba
|
6
|
--
|
--
|
--
|
--
|
| |
DeepHermes 3 - Mistral 24B
|
32k
|
Nous Research
|
5
*
|
--
|
--
|
--
|
--
|
| |
Jamba 1.7 Large
|
256k
|
AI21 Labs
|
5
*
|
--
|
54
|
1.40
|
10.60
|
| |
Granite 4.0 H Small
|
128k
|
IBM
|
5
*
|
--
|
407
|
10.22
|
11.45
|
| |
Qwen3 Omni 30B A3B
|
65.5k
|
Alibaba
|
5
*
|
--
|
94
|
1.89
|
7.19
|
| |
LFM2 24B A2B
|
32.8k
|
Liquid AI
|
5
*
|
--
|
--
|
--
|
--
|
| |
Phi-4
|
16k
|
Microsoft
|
5
*
|
--
|
--
|
--
|
--
|
| |
Nova Micro
|
130k
|
Amazon
|
5
*
|
--
|
298
|
0.90
|
2.58
|
| |
Granite 4.1 3B
|
131k
|
IBM
|
5
|
--
|
--
|
--
|
--
|
| |
NVIDIA Nemotron Nano 12B v2 VL
|
128k
|
NVIDIA
|
5
*
|
--
|
172
|
2.21
|
5.12
|
| |
Phi-4 Multimodal
|
128k
|
Microsoft
|
5
*
|
--
|
18
|
0.85
|
28.48
|
| |
MiniCPM-V 4.6 1.3B
|
262k
|
OpenBMB
|
4
|
--
|
--
|
--
|
--
|
| |
Jamba Reasoning 3B
|
262k
|
AI21 Labs
|
4
*
|
--
|
--
|
--
|
--
|
| |
Reka Flash 3
|
128k
|
Reka AI
|
4
*
|
--
|
--
|
--
|
--
|
| |
Olmo 3 7B Think
|
65.5k
|
Allen Institute for AI
|
4
*
|
--
|
--
|
--
|
--
|
| |
Molmo 7B-D
|
4.1k
|
Allen Institute for AI
|
4
*
|
--
|
--
|
--
|
--
|
| |
Ling-mini-2.0
|
131k
|
InclusionAI
|
4
*
|
--
|
--
|
--
|
--
|
| |
Llama 3.2 11B (Vision)
|
128k
|
Meta
|
3
*
|
--
|
8
|
2.59
|
68.73
|
| |
Qwen3.5 0.8B
|
262k
|
Alibaba
|
3
|
--
|
--
|
--
|
--
|
| |
Exaone 4.0 1.2B
|
64k
|
LG AI Research
|
3
*
|
--
|
--
|
--
|
--
|
| |
Olmo 3 7B
|
65.5k
|
Allen Institute for AI
|
3
*
|
--
|
--
|
--
|
--
|
| |
Exaone 4.0 1.2B
|
64k
|
LG AI Research
|
3
*
|
--
|
--
|
--
|
--
|
| |
LFM2.5-1.2B-Thinking
|
32k
|
Liquid AI
|
3
*
|
--
|
--
|
--
|
--
|
| |
Jamba 1.7 Mini
|
258k
|
AI21 Labs
|
3
*
|
--
|
--
|
--
|
--
|
| |
LFM2 2.6B
|
32.8k
|
Liquid AI
|
3
*
|
--
|
--
|
--
|
--
|
| |
LFM2.5-1.2B-Instruct
|
32k
|
Liquid AI
|
3
*
|
--
|
--
|
--
|
--
|
| |
Granite 4.0 H 1B
|
128k
|
IBM
|
3
*
|
--
|
--
|
--
|
--
|
| |
Gemma 3 270M
|
32k
|
Google
|
2
*
|
--
|
--
|
--
|
--
|
| |
Apertus 70B Instruct
|
65.5k
|
Swiss AI Initiative
|
2
*
|
--
|
--
|
--
|
--
|
| |
Granite 4.0 Micro
|
128k
|
IBM
|
2
*
|
--
|
--
|
--
|
--
|
| |
DeepHermes 3 - Llama-3.1 8B
|
128k
|
Nous Research
|
2
*
|
--
|
--
|
--
|
--
|
| |
Granite 4.0 1B
|
128k
|
IBM
|
2
*
|
--
|
--
|
--
|
--
|
| |
Molmo2-8B
|
36.9k
|
Allen Institute for AI
|
2
*
|
--
|
--
|
--
|
--
|
| |
LFM2 8B A1B
|
32.8k
|
Liquid AI
|
2
*
|
--
|
--
|
--
|
--
|
| |
LFM2.5-VL-1.6B
|
32k
|
Liquid AI
|
1
*
|
--
|
369
|
11.90
|
13.25
|
| |
Granite 4.0 350M
|
32.8k
|
IBM
|
1
*
|
--
|
--
|
--
|
--
|
| |
Tiny Aya Global
|
8.19k
|
Cohere
|
1
*
|
--
|
--
|
--
|
--
|
| |
Apertus 8B Instruct
|
65.5k
|
Swiss AI Initiative
|
1
*
|
--
|
--
|
--
|
--
|
| |
Granite 4.0 H 350M
|
32.8k
|
IBM
|
1
*
|
--
|
--
|
--
|
--
|
| |
Claude Sonnet 5 (low)
|
1M
|
Anthropic
|
--
|
--
|
63
|
1.65
|
9.64
|
| |
EXAONE 4.5 33B
|
262k
|
LG AI Research
|
--
|
--
|
--
|
--
|
--
|
| |
Claude Sonnet 5 (xhigh)
|
1M
|
Anthropic
|
--
|
--
|
66
|
29.93
|
37.47
|
| |
Gemini 3 Deep Think
|
128k
|
Google
|
--
|
--
|
--
|
--
|
--
|
| |
Claude Sonnet 5 (high)
|
1M
|
Anthropic
|
--
|
--
|
63
|
14.23
|
22.16
|
| |
Mi:dm K 2.5 Pro Preview
|
128k
|
Korea Telecom
|
--
|
--
|
--
|
--
|
--
|
| |
GPT-5.5 Pro (xhigh)
|
922k
|
OpenAI
|
--
|
--
|
--
|
--
|
--
|
| |
Cogito v2.1
|
128k
|
Deep Cogito
|
--
|
--
|
--
|
--
|
--
|
| |
Claude Sonnet 5 (medium)
|
1M
|
Anthropic
|
--
|
--
|
63
|
1.89
|
9.78
|
|
Maximum number of combined input & output tokens. Output tokens commonly have a significantly lower limit (varied by model).
Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).
Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.
Average cost per task in the index. Costs are split by input, cache hit, cache write, reasoning, and answer token pricing where canonical token counts are available.
Price per token included in the request/message sent to the API, represented as USD per million Tokens.
Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.
Price per token to write prompt tokens into the cache so that later requests can hit them, represented as USD per million tokens. Some providers charge a premium over the standard input price to create a cache entry (e.g. Anthropic), while others cache automatically with no separate write fee.
Price per token generated by the model (received from the API), represented as USD per million Tokens.
Metrics are 'live' and are based on the past 72 hours of measurements, measurements are taken 8 times a day for single requests and 2 times per day for parallel requests.
Claude Opus 5 (Adaptive Reasoning, Max Effort) currently ranks #1 on the Artificial Analysis LLM Leaderboard with an Intelligence Index score of 61, out of 170 models ranked.
The top models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).
Mercury 2 is the fastest at 901.6 tokens per second, followed by HyperNova 60B 2605 (414.2 t/s) and Granite 4.0 H Small (407.1 t/s).
Llama 4 Scout has the lowest cost per Intelligence Index task at $0.01, followed by MiMo-V2.5 ($0.01) and gpt-oss-120b (low) ($0.02).
GLM-5.2 (max) is the highest-ranked open weights model with an Intelligence Index score of 51. There are 94 open weights models out of 170 total on the leaderboard.
The top open weights models by Intelligence Index are: 1. GLM-5.2 (max) (51), 2. MiniMax-M3 (44), 3. DeepSeek V4 Pro (Reasoning, Max Effort) (44).
Claude Opus 5 (Adaptive Reasoning, Max Effort) leads among 126 reasoning models with an Intelligence Index score of 61. Reasoning models use extended thinking to solve complex problems before responding.
The leaderboard includes filters to narrow results by model type (reasoning vs non-reasoning), openness (open weights vs proprietary), and other criteria. You can also adjust prompt options to see how performance varies with different input lengths.
Click on any model name in the leaderboard to visit its dedicated comparison page with detailed charts covering intelligence, pricing, speed, latency, and more. You can also compare API providers for each model. View all models