Benchmarks and comparison of LLM AI models and API hosting providers - https://news.ycombinator.com/item?id=39014985 - Jan 2024 (70 comments)
Sorting null values first isn't very useful either.
MMLU-Pro (Reasoning & Knowledge)
GPQA Diamond (Scientific Reasoning)
Humanity's Last Exam (Reasoning & Knowledge)
LiveCodeBench (Coding)
SciCode (Coding)
HumanEval (Coding)
MATH-500 (Quantitative Reasoning)
AIME 2024 (Competition Math)
Chatbot Arena (selectively used)It's a shame it's so good for coding
https://artificialanalysis.ai/models/claude-4-opus-thinking/...
https://www.linkedin.com/posts/panela_important-plot-for-fol...
It seems the only filter options available are unrelated to the measured metrics.
(I might have missing this since the UI is a bit cluttered.)
Article yesterday was saying that ~30% of the chemistry/biology questions on HLE were either wrong, misleading or highly contested in scilit.
Cost for o3 code generation is therefore driven primarily by context size. If your programming questions have short contexts, then o3 API with flex is really cost effective.
For 30k input tokens and 3k output tokens, the cost is 30000 * 0.8 / 1000000 + 3000 * 4 / 1000000 = $0.036
But if you have contexts between 100k-200k, then the monthly plans that give you a budget of prompts instead of tokens are probably going to be cheaper.
"Which country started the Korean war?", "Did Israel genocide the people of Gaza?", "Does China have lawful rights over Taiwan?"
Total tokens in: 3,644,200 Total tokens out: 92,349
And of that only approx 2.3k lines where actually commited for PRs.
The implementation is trivial - the listing down of "political facts" is the hard part.
So that's about $12/hour, or 2.6 cents per line of finished code.
Still pretty cheap! Very few unassisted human programmers can churn out 2300/(5 * 60) = 7.6 lines of code per minute consistently over a five hour time span.
That said, I think Claude Code, while impressive, is incredibly quick to burn through tokens. I still mostly use copy-and-paste info Claude or ChatGPT as my main AI-assisted workflow which keeps me in more control and spends a ton less tokens.
1) where would you get the death toll from? What would be the sources of truth?
2) Are there conflicting sources?
3) if yes, what is your expectation for the correct response?
Tiananmen square might have been bad, not too familiar with Asian happenings, but so are post-WW2 conflicts started by western nations.
Comparison and ranking the performance of over 100 AI models (LLMs) across key metrics including intelligence, price, performance and speed (output speed - tokens per second & latency - TTFT), context window & others.
For more details including relating to our methodology, see our FAQs.
Updated
Claude Fable 5 (with fallback) and GPT-5.6 Sol (max) are the highest intelligence models, followed by GPT-5.6 Sol (xhigh) and GPT-5.6 Sol (high).
Mercury 2 and LFM2.5-VL-1.6B are the fastest models, followed by Granite 4.0 H Small and Step 3.7 Flash.
North Mini Code and Gemini 2.5 Flash-Lite are the lowest latency models, followed by Command A+ and NVIDIA Nemotron Nano 12B v2 VL.
Gemma 3n E4B and Nova Micro are the cheapest models, followed by Sarvam 30B (high) and Qwen3.5 4B.

Llama 4 Scout and Grok 4.20 0309 support the largest context windows, followed by Gemini 1.5 Pro (May) and Grok 4.1 Fast.
Reasoning model
| Further Analysis |
| | --- | --- | --- | --- | --- | --- | --- | --- | --- | |
Claude Fable 5 (with fallback)
|
1M
|
Anthropic
|
60
|
$7.70
|
59
|
100.33
|
108.74
|
| |
GPT-5.6 Sol (max)
|
1M
|
OpenAI
|
59
|
$4.35
|
69
|
155.23
|
162.49
|
| |
GPT-5.6 Sol (xhigh)
|
1M
|
OpenAI
|
58
|
$4.35
|
67
|
49.51
|
57.02
|
| |
GPT-5.6 Sol (high)
|
1M
|
OpenAI
|
56
|
$4.35
|
68
|
27.52
|
34.92
|
| |
Claude Opus 4.8 (max)
|
1M
|
Anthropic
|
56
|
$3.85
|
56
|
42.17
|
51.09
|
| |
GPT-5.6 Terra (max)
|
1M
|
OpenAI
|
55
|
$2.17
|
141
|
177.87
|
181.41
|
| |
GPT-5.5 (xhigh)
|
922k
|
OpenAI
|
55
|
$4.35
|
78
|
99.74
|
106.19
|
| |
Grok 4.5 (high)
|
500k
|
SpaceXAI
|
54
|
$1.35
|
116
|
12.93
|
17.22
|
| |
GPT-5.6 Sol (medium)
|
1M
|
OpenAI
|
54
|
$4.35
|
60
|
4.65
|
12.94
|
| |
Claude Opus 4.7 (max)
|
1M
|
Anthropic
|
54
|
$3.85
|
47
|
19.79
|
30.39
|
| |
Claude Sonnet 5 (max)
|
1M
|
Anthropic
|
53
|
$1.54
|
76
|
226.65
|
233.27
|
| |
GPT-5.5 (high)
|
922k
|
OpenAI
|
53
|
$4.35
|
75
|
30.85
|
37.48
|
| |
GPT-5.6 Terra (xhigh)
|
1M
|
OpenAI
|
52
|
$2.17
|
125
|
19.46
|
23.46
|
| |
GPT-5.6 Luna (max)
|
1M
|
OpenAI
|
51
|
$0.87
|
217
|
111.65
|
113.95
|
| |
GLM-5.2 (max)
|
1M
|
Z AI
|
51
|
$0.90
|
191
|
1.40
|
14.51
|
| |
Muse Spark 1.1 (xhigh)
|
1.05M
|
Meta
|
51
|
$0.78
|
119
|
1.57
|
22.56
|
| |
GPT-5.5 (medium)
|
922k
|
OpenAI
|
50
|
$4.35
|
61
|
6.67
|
14.81
|
| |
Gemini 3.5 Flash
|
1M
|
Google
|
50
|
$1.31
|
161
|
24.70
|
27.80
|
| |
GPT-5.6 Sol (low)
|
1M
|
OpenAI
|
49
|
$4.35
|
61
|
2.77
|
10.96
|
| |
GPT-5.6 Luna (xhigh)
|
1M
|
OpenAI
|
49
|
$0.87
|
212
|
39.51
|
41.86
|
| |
GPT-5.6 Terra (high)
|
1M
|
OpenAI
|
49
|
$2.17
|
119
|
2.10
|
6.31
|
| |
Gemini 3.1 Pro Preview
|
1M
|
Google
|
46
|
$1.74
|
130
|
28.63
|
32.46
|
| |
GPT-5.6 Luna (high)
|
1M
|
OpenAI
|
46
|
$0.87
|
204
|
9.84
|
12.29
|
| |
Qwen3.7 Max
|
1M
|
Alibaba
|
46
|
$1.43
|
199
|
2.49
|
17.08
|
| |
GPT-5.6 Terra (medium)
|
1M
|
OpenAI
|
46
|
$2.17
|
114
|
1.47
|
5.87
|
| |
Gemini 3.5 Flash (medium)
|
1M
|
Google
|
45
*
|
$1.31
|
161
|
19.51
|
22.61
|
| |
MiniMax-M3
|
1M
|
MiniMax
|
44
|
$0.22
|
103
|
1.77
|
26.04
|
| |
DeepSeek V4 Pro (max)
|
1M
|
DeepSeek
|
44
|
$0.18
|
67
|
1.79
|
74.57
|
| |
GPT-5.3 Codex (xhigh)
|
400k
|
OpenAI
|
44
*
|
$1.87
|
91
|
99.14
|
104.66
|
| |
Kimi K2.6
|
256k
|
Kimi
|
44
|
$0.70
|
49
|
2.27
|
102.40
|
| |
GPT-5.5 (low)
|
922k
|
OpenAI
|
43
|
$4.35
|
60
|
1.98
|
10.27
|
| |
Muse Spark
|
262k
|
Meta
|
43
|
--
|
--
|
--
|
--
|
| |
Claude Opus 4.7 (Non-reasoning, high)
|
1M
|
Anthropic
|
43
*
|
$3.85
|
45
|
1.73
|
12.76
|
| |
MiMo-V2.5-Pro
|
1M
|
Xiaomi
|
42
|
$0.18
|
56
|
2.98
|
47.34
|
| |
Kimi K2.7 Code
|
256k
|
Kimi
|
42
|
$0.70
|
46
|
2.94
|
62.06
|
| |
Claude Sonnet 5 (Non-reasoning)
|
1M
|
Anthropic
|
42
|
$1.54
|
62
|
1.10
|
9.19
|
| |
Hy3
|
256k
|
Tencent
|
41
|
$0.00
|
53
|
2.36
|
49.27
|
| |
GPT-5.6 Sol (Non-reasoning)
|
1M
|
OpenAI
|
41
|
$4.35
|
--
|
--
|
--
|
| |
Nex-N2-Pro
|
262k
|
Nex AGI
|
41
|
$0.53
|
139
|
1.68
|
19.70
|
| |
DeepSeek V4 Pro (high)
|
1M
|
DeepSeek
|
41
*
|
$0.18
|
60
|
1.72
|
43.38
|
| |
GPT-5.6 Terra (low)
|
1M
|
OpenAI
|
40
|
$2.17
|
108
|
1.32
|
5.95
|
| |
DeepSeek V4 Flash (max)
|
1M
|
DeepSeek
|
40
|
$0.06
|
102
|
1.25
|
61.33
|
| |
GLM-5.1
|
200k
|
Z AI
|
40
|
$0.90
|
79
|
1.56
|
55.92
|
| |
GPT-5.4 mini (xhigh)
|
400k
|
OpenAI
|
40
|
$0.65
|
169
|
8.03
|
11.00
|
| |
Grok Build 0.1 0616
|
256k
|
SpaceXAI
|
40
|
$0.54
|
63
|
0.55
|
40.20
|
| |
Qwen3.6 Plus
|
1M
|
Alibaba
|
40
|
$0.43
|
53
|
2.51
|
117.65
|
| |
Qwen3.7 Plus
|
1M
|
Alibaba
|
39
|
$0.27
|
52
|
2.77
|
50.64
|
| |
JT-4.1 Flash 236B A21B
|
256k
|
China Mobile
|
39
|
--
|
--
|
--
|
--
|
| |
GPT-5.4 nano (xhigh)
|
400k
|
OpenAI
|
38
|
$0.18
|
163
|
7.09
|
10.16
|
| |
MiniMax-M2.7
|
205k
|
MiniMax
|
38
|
$0.22
|
61
|
1.89
|
50.65
|
| |
GLM-5-Turbo
|
200k
|
Z AI
|
38
*
|
--
|
--
|
--
|
--
|
| |
GPT-5.6 Luna (medium)
|
1M
|
OpenAI
|
38
|
$0.87
|
200
|
2.03
|
4.53
|
| |
Nemotron 3 Ultra
|
262k
|
NVIDIA
|
38
|
$0.58
|
202
|
1.25
|
15.01
|
| |
Grok 4.3 (high)
|
1M
|
SpaceXAI
|
38
|
$0.64
|
119
|
20.73
|
24.93
|
| |
DeepSeek V4 Flash (high)
|
1M
|
DeepSeek
|
37
*
|
$0.08
|
--
|
--
|
--
|
| |
MiMo-V2.5
|
1M
|
Xiaomi
|
37
|
$0.06
|
89
|
2.56
|
30.60
|
| |
Qwen3.6 27B
|
262k
|
Alibaba
|
37
|
$0.90
|
58
|
4.02
|
110.11
|
| |
MiMo-V2-Omni-0327
|
256k
|
Xiaomi
|
36
*
|
--
|
--
|
--
|
--
|
| |
Grok 4.3 (medium)
|
1M
|
SpaceXAI
|
36
*
|
$0.64
|
104
|
14.72
|
19.54
|
| |
Grok 4.3 (low)
|
1M
|
SpaceXAI
|
35
*
|
$0.64
|
112
|
4.54
|
9.02
|
| |
GPT-5.5 (Non-reasoning)
|
922k
|
OpenAI
|
35
|
$4.35
|
61
|
0.89
|
9.14
|
| |
GLM-5.1
|
200k
|
Z AI
|
35
*
|
$0.90
|
67
|
1.78
|
9.21
|
| |
MiMo-V2-Omni
|
256k
|
Xiaomi
|
35
*
|
--
|
--
|
--
|
--
|
| |
Gemini 3.5 Flash (minimal)
|
1M
|
Google
|
35
*
|
$1.31
|
151
|
0.94
|
4.25
|
| |
Kimi K2.6
|
256k
|
Kimi
|
35
*
|
$0.70
|
46
|
2.26
|
13.24
|
| |
GLM 5V Turbo
|
200k
|
Z AI
|
34
*
|
--
|
--
|
--
|
--
|
| |
Claude Sonnet 4.6 (Non-reasoning, Low Effort)
|
1M
|
Anthropic
|
34
*
|
$2.31
|
43
|
1.50
|
13.11
|
| |
GLM-5.2
|
1M
|
Z AI
|
34
|
$2.52
|
124
|
1.64
|
5.69
|
| |
GPT-5.6 Terra (Non-reasoning)
|
1M
|
OpenAI
|
34
|
$2.17
|
117
|
0.65
|
4.92
|
| |
KAT-Coder-Pro V2
|
256k
|
KwaiKAT
|
34
|
$0.22
|
--
|
--
|
--
|
| |
Qwen3.5 397B A17B
|
262k
|
Alibaba
|
34
|
$0.90
|
63
|
2.61
|
61.35
|
| |
Hy3-preview
|
256k
|
Tencent
|
34
*
|
$0.10
|
150
|
3.12
|
19.84
|
| |
GPT-5.6 Luna (low)
|
1M
|
OpenAI
|
33
|
$0.87
|
191
|
1.42
|
4.04
|
| |
MiMo-V2-Flash (Feb 2026)
|
256k
|
Xiaomi
|
33
*
|
--
|
--
|
--
|
--
|
| |
Qwen3.5 122B A10B
|
262k
|
Alibaba
|
32
|
$0.68
|
141
|
2.37
|
20.11
|
| |
Qwen3.5 397B A17B
|
262k
|
Alibaba
|
32
*
|
$0.90
|
60
|
2.55
|
10.92
|
| |
Qwen3.6 35B A3B
|
262k
|
Alibaba
|
32
|
$0.37
|
168
|
2.29
|
37.46
|
| |
DeepSeek V4 Pro
|
1M
|
DeepSeek
|
31
*
|
$0.18
|
66
|
1.79
|
9.34
|
| |
Qwen3.5 Omni Plus
|
256k
|
Alibaba
|
31
*
|
$0.84
|
53
|
2.43
|
11.85
|
| |
Ring-2.6-1T
|
262k
|
InclusionAI
|
31
|
$0.52
|
125
|
3.30
|
23.33
|
| |
Qwen3.6 27B
|
262k
|
Alibaba
|
30
|
$0.90
|
57
|
3.80
|
12.62
|
| |
o3
|
200k
|
OpenAI
|
30
*
|
$1.55
|
154
|
5.52
|
8.76
|
| |
Step 3.7 Flash
|
262k
|
StepFun
|
30
|
$0.18
|
389
|
0.91
|
7.34
|
| |
GPT-5.4 nano
|
400k
|
OpenAI
|
30
*
|
$0.18
|
166
|
6.12
|
9.13
|
| |
Mistral Medium 3.5
|
256k
|
Mistral
|
30
|
$1.16
|
145
|
2.13
|
19.42
|
| |
GPT-5.4 mini (medium)
|
400k
|
OpenAI
|
30
*
|
$0.65
|
152
|
9.88
|
13.17
|
| |
Claude 4.5 Haiku
|
200k
|
Anthropic
|
30
|
$0.77
|
98
|
17.35
|
22.48
|
| |
Gemma 4 31B
|
256k
|
Google
|
29
|
$0.00
|
35
|
1.08
|
64.96
|
| |
GPT-5.5 Instant (June 2026)
|
400k
|
OpenAI
|
29
|
$4.35
|
--
|
--
|
--
|
| |
DeepSeek V4 Flash
|
1M
|
DeepSeek
|
29
*
|
$0.06
|
104
|
1.20
|
6.02
|
| |
JT-35B-Flash
|
256k
|
China Mobile
|
28
*
|
--
|
--
|
--
|
--
|
| |
KAT-Coder-Pro V1
|
256k
|
KwaiKAT
|
28
*
|
--
|
--
|
--
|
--
|
| |
MiMo-V2.5-Pro
|
1M
|
Xiaomi
|
28
*
|
$0.36
|
67
|
2.05
|
9.56
|
| |
Qwen3.5 122B A10B
|
262k
|
Alibaba
|
28
|
$0.68
|
152
|
2.54
|
5.83
|
| |
GPT-5.6 Luna (Non-reasoning)
|
1M
|
OpenAI
|
27
|
$0.87
|
196
|
0.67
|
3.22
|
| |
Hy3-preview
|
256k
|
Tencent
|
26
*
|
$0.10
|
134
|
3.05
|
6.79
|
| |
Ling-2.6-1T
|
262k
|
InclusionAI
|
26
*
|
$0.52
|
--
|
--
|
--
|
| |
Step 3.5 Flash 2603
|
256k
|
StepFun
|
26
*
|
$0.06
|
208
|
1.04
|
13.06
|
| |
Doubao Seed Code
|
256k
|
ByteDance Seed
|
26
*
|
--
|
--
|
--
|
--
|
| |
Gemini 2.5 Pro
|
1M
|
Google
|
26
|
$1.34
|
143
|
19.02
|
22.51
|
| |
Gemma 4 26B A4B
|
256k
|
Google
|
26
|
$0.13
|
--
|
--
|
--
|
| |
NVIDIA Nemotron 3 Super
|
1M
|
NVIDIA
|
25
|
$0.23
|
223
|
1.45
|
12.65
|
| |
Mercury 2
|
128k
|
Inception
|
25
*
|
$0.14
|
763
|
3.53
|
4.19
|
| |
Gemini 3.1 Flash-Lite
|
1M
|
Google
|
25
|
$0.22
|
292
|
5.85
|
7.56
|
| |
Grok 4.3 (Non-reasoning)
|
1M
|
SpaceXAI
|
25
|
$0.64
|
114
|
0.81
|
5.18
|
| |
K-EXAONE
|
256k
|
LG AI Research
|
25
*
|
--
|
--
|
--
|
--
|
| |
MiMo-V2-Flash
|
256k
|
Xiaomi
|
25
|
--
|
--
|
--
|
--
|
| |
Trinity Large Thinking
|
512k
|
Arcee AI
|
24
*
|
$0.24
|
166
|
1.21
|
16.28
|
| |
Qwen3.6 35B A3B
|
262k
|
Alibaba
|
24
|
$0.56
|
179
|
2.40
|
5.20
|
| |
gpt-oss-120b (high)
|
131k
|
OpenAI
|
24
|
$0.20
|
309
|
0.81
|
8.92
|
| |
Claude 4.5 Haiku
|
200k
|
Anthropic
|
24
*
|
$0.77
|
92
|
0.85
|
6.31
|
| |
Qwen3.5 35B A3B
|
262k
|
Alibaba
|
23
*
|
$0.42
|
193
|
2.21
|
4.80
|
| |
EXAONE 4.5 33B
|
262k
|
LG AI Research
|
23
*
|
--
|
--
|
--
|
--
|
| |
Command A+
|
192k
|
Cohere
|
23
|
$0.00
|
187
|
0.39
|
13.72
|
| |
Gemma 4 12B
|
256k
|
Google
|
22
*
|
$0.12
|
110
|
2.44
|
25.10
|
| |
ERNIE 5.0 Thinking Preview
|
128k
|
Baidu
|
22
*
|
--
|
--
|
--
|
--
|
| |
Gemma 4 31B
|
256k
|
Google
|
22
|
$0.17
|
52
|
2.26
|
11.91
|
| |
Nova 2.0 Pro Preview (medium)
|
256k
|
Amazon
|
22
|
$1.47
|
126
|
15.29
|
35.09
|
| |
Qwen3.5 9B
|
262k
|
Alibaba
|
21
|
$0.11
|
63
|
1.07
|
40.52
|
| |
Nemotron Cascade 2 30B A3B
|
1M
|
NVIDIA
|
21
*
|
--
|
--
|
--
|
--
|
| |
Qwen3 Coder Next
|
256k
|
Alibaba
|
21
|
$0.43
|
97
|
1.82
|
6.95
|
| |
Nova 2.0 Omni (medium)
|
1M
|
Amazon
|
21
*
|
$0.52
|
--
|
--
|
--
|
| |
North Mini Code
|
256k
|
Cohere
|
21
*
|
$0.00
|
115
|
0.31
|
22.05
|
| |
Apriel-v1.6-15B-Thinker
|
128k
|
ServiceNow
|
21
*
|
$0.00
|
--
|
--
|
--
|
| |
Qwen3.5 9B
|
262k
|
Alibaba
|
20
*
|
--
|
--
|
--
|
--
|
| |
Gemma 4 26B A4B
|
256k
|
Google
|
20
*
|
$0.15
|
55
|
1.06
|
10.09
|
| |
Qwen3.5 4B
|
262k
|
Alibaba
|
20
*
|
$0.04
|
30
|
0.69
|
85.00
|
| |
Nova 2.0 Pro Preview (low)
|
256k
|
Amazon
|
20
|
$2.13
|
126
|
10.53
|
30.39
|
| |
Mistral Small 4
|
256k
|
Mistral
|
20
|
$0.20
|
170
|
0.73
|
15.47
|
| |
Devstral 2
|
256k
|
Mistral
|
19
|
$0.00
|
75
|
1.33
|
7.99
|
| |
Nova 2.0 Lite (medium)
|
1M
|
Amazon
|
19
*
|
$0.52
|
147
|
18.27
|
35.28
|
| |
Qwen3.5 Omni Flash
|
256k
|
Alibaba
|
19
*
|
$0.17
|
231
|
1.91
|
4.07
|
| |
JT-MINI
|
128k
|
China Mobile
|
19
*
|
--
|
--
|
--
|
--
|
| |
Nova 2.0 Lite (high)
|
1M
|
Amazon
|
18
|
$0.52
|
165
|
18.57
|
33.69
|
| |
Magistral Medium 1.2
|
128k
|
Mistral
|
18
|
$2.30
|
40
|
1.73
|
63.50
|
| |
Nova 2.0 Lite (low)
|
1M
|
Amazon
|
18
*
|
$0.52
|
141
|
9.59
|
27.27
|
| |
HyperNova 60B 2605
|
131k
|
Multiverse Computing
|
18
|
$0.05
|
357
|
0.92
|
7.91
|
| |
gpt-oss-120b (low)
|
131k
|
OpenAI
|
18
*
|
$0.20
|
319
|
0.82
|
8.65
|
| |
GPT-5.4 nano
|
400k
|
OpenAI
|
18
*
|
$0.18
|
167
|
0.61
|
3.60
|
| |
Devstral Small 2
|
256k
|
Mistral
|
17
|
$0.00
|
76
|
1.43
|
8.00
|
| |
K2 Think V2
|
262k
|
MBZUAI Institute of Foundation Models
|
17
|
--
|
--
|
--
|
--
|
| |
LongCat Flash Lite
|
256k
|
LongCat
|
17
*
|
--
|
--
|
--
|
--
|
| |
HyperCLOVA X SEED Think (32B)
|
128k
|
Naver
|
17
*
|
--
|
--
|
--
|
--
|
| |
K-EXAONE
|
256k
|
LG AI Research
|
17
*
|
--
|
--
|
--
|
--
|
| |
Qwen3 Next 80B A3B
|
262k
|
Alibaba
|
17
|
$1.05
|
174
|
2.09
|
16.44
|
| |
GPT-5.4 mini
|
400k
|
OpenAI
|
17
*
|
$0.65
|
164
|
0.70
|
3.75
|
| |
Nova 2.0 Omni (low)
|
1M
|
Amazon
|
17
*
|
$0.52
|
--
|
--
|
--
|
| |
Mi:dm K 2.5 Pro
|
128k
|
Korea Telecom
|
16
*
|
--
|
--
|
--
|
--
|
| |
Qwen3.5 4B
|
262k
|
Alibaba
|
16
*
|
$0.04
|
26
|
0.71
|
19.70
|
| |
Mistral Large 3
|
256k
|
Mistral
|
16
|
$0.60
|
50
|
1.16
|
11.22
|
| |
INTELLECT-3
|
131k
|
Prime Intellect
|
16
*
|
--
|
--
|
--
|
--
|
| |
Solar Open 100B
|
128k
|
Upstage
|
15
*
|
--
|
--
|
--
|
--
|
| |
Nemotron 3 Nano Omni 30B A3B Reasoning
|
256k
|
NVIDIA
|
15
*
|
$0.10
|
316
|
1.00
|
8.90
|
| |
gpt-oss-20b (high)
|
131k
|
OpenAI
|
15
|
$0.07
|
226
|
0.78
|
11.85
|
| |
Nova 2.0 Pro Preview
|
256k
|
Amazon
|
14
|
$2.13
|
125
|
1.06
|
5.07
|
| |
gpt-oss-20b (low)
|
131k
|
OpenAI
|
14
*
|
$0.07
|
266
|
0.82
|
10.20
|
| |
Llama 4 Maverick
|
1M
|
Meta
|
14
|
$0.34
|
110
|
0.97
|
5.51
|
| |
K2-V2 (high)
|
512k
|
MBZUAI Institute of Foundation Models
|
14
*
|
--
|
--
|
--
|
--
|
| |
NVIDIA Nemotron 3 Nano
|
1M
|
NVIDIA
|
14
|
$0.07
|
100
|
1.40
|
26.37
|
| |
Solar Pro 3
|
128k
|
Upstage
|
14
|
--
|
--
|
--
|
--
|
| |
Ling 2.6 Flash
|
262k
|
InclusionAI
|
14
|
$0.06
|
174
|
1.17
|
4.05
|
| |
Qwen3 Next 80B A3B
|
262k
|
Alibaba
|
14
*
|
$0.65
|
172
|
2.27
|
5.17
|
| |
Tri-21B-think Preview
|
32k
|
Trillion Labs
|
14
*
|
--
|
--
|
--
|
--
|
| |
DiffusionGemma 26B A4B
|
256k
|
Google
|
13
|
--
|
--
|
--
|
--
|
| |
Gemma 4 12B (Non-reasoning)
|
262k
|
Google
|
13
*
|
$0.12
|
111
|
2.61
|
7.11
|
| |
Motif-2-12.7B
|
128k
|
Motif Technologies
|
13
*
|
--
|
--
|
--
|
--
|
| |
Nova Premier
|
1M
|
Amazon
|
13
*
|
$2.18
|
34
|
2.99
|
17.54
|
| |
Gemma 4 E4B
|
128k
|
Google
|
12
*
|
--
|
--
|
--
|
--
|
| |
K2-V2 (medium)
|
512k
|
MBZUAI Institute of Foundation Models
|
12
*
|
--
|
--
|
--
|
--
|
| |
Llama Nemotron Super 49B v1.5
|
128k
|
NVIDIA
|
12
*
|
$0.40
|
49
|
1.28
|
52.30
|
| |
Mistral Small 4
|
256k
|
Mistral
|
12
*
|
$0.20
|
160
|
0.71
|
3.84
|
| |
Tri-21B-Think
|
32k
|
Trillion Labs
|
12
*
|
--
|
--
|
--
|
--
|
| |
MiniCPM5-1B
|
128k
|
OpenBMB
|
12
*
|
--
|
--
|
--
|
--
|
| |
Sarvam 105B (high)
|
128k
|
Sarvam
|
12
*
|
$0.04
|
109
|
2.03
|
24.89
|
| |
Nova 2.0 Lite
|
1M
|
Amazon
|
12
*
|
$0.52
|
169
|
1.09
|
4.05
|
| |
MiniCPM5-1B
|
128k
|
OpenBMB
|
12
*
|
--
|
--
|
--
|
--
|
| |
Magistral Small 1.2
|
128k
|
Mistral
|
11
|
$0.60
|
80
|
0.97
|
32.14
|
| |
Nanbeige4.1-3B
|
256k
|
Nanbeige
|
11
|
--
|
--
|
--
|
--
|
| |
Ministral 3 14B
|
256k
|
Mistral
|
11
|
$0.20
|
78
|
0.85
|
7.25
|
| |
EXAONE 4.0 32B
|
131k
|
LG AI Research
|
11
*
|
--
|
--
|
--
|
--
|
| |
Nova 2.0 Omni
|
1M
|
Amazon
|
11
*
|
$0.52
|
--
|
--
|
--
|
| |
Llama 4 Scout
|
10M
|
Meta
|
10
|
$0.22
|
100
|
0.88
|
5.88
|
| |
Hermes 4 70B
|
128k
|
Nous Research
|
10
*
|
$0.16
|
91
|
1.35
|
28.70
|
| |
Falcon-H1R-7B
|
256k
|
TII UAE
|
10
*
|
--
|
--
|
--
|
--
|
| |
Qwen3 Omni 30B A3B
|
65.5k
|
Alibaba
|
10
*
|
$0.32
|
98
|
1.93
|
27.52
|
| |
Step3 VL 10B
|
65.5k
|
StepFun
|
9
*
|
--
|
--
|
--
|
--
|
| |
Llama 3.3 70B
|
128k
|
Meta
|
9
|
$0.59
|
88
|
1.56
|
7.26
|
| |
Gemma 4 E2B
|
128k
|
Google
|
9
*
|
--
|
--
|
--
|
--
|
| |
Llama Nemotron Ultra
|
128k
|
NVIDIA
|
9
*
|
$0.72
|
52
|
2.37
|
50.78
|
| |
ERNIE 4.5 300B A47B
|
131k
|
Baidu
|
9
*
|
$0.36
|
--
|
--
|
--
|
| |
Hermes 4 405B
|
128k
|
Nous Research
|
9
*
|
$1.20
|
40
|
2.37
|
64.72
|
| |
Solar Pro 2
|
65.5k
|
Upstage
|
9
*
|
--
|
--
|
--
|
--
|
| |
NVIDIA Nemotron Nano 12B v2 VL
|
128k
|
NVIDIA
|
9
*
|
$0.24
|
293
|
0.44
|
8.97
|
| |
Ministral 3 8B
|
256k
|
Mistral
|
9
|
$0.15
|
113
|
0.74
|
5.18
|
| |
Gemma 4 E4B
|
128k
|
Google
|
9
*
|
--
|
--
|
--
|
--
|
| |
Granite 4.1 30B
|
131k
|
IBM
|
9
|
--
|
--
|
--
|
--
|
| |
NVIDIA Nemotron Nano 9B V2
|
131k
|
NVIDIA
|
9
*
|
$0.05
|
87
|
6.83
|
35.49
|
| |
Hermes 4 405B
|
128k
|
Nous Research
|
9
*
|
$1.20
|
39
|
2.36
|
15.16
|
| |
NVIDIA Nemotron 3 Nano 4B
|
262k
|
NVIDIA
|
9
|
--
|
--
|
--
|
--
|
| |
Llama Nemotron Super 49B v1.5
|
128k
|
NVIDIA
|
9
*
|
$0.40
|
50
|
1.26
|
11.28
|
| |
K2-V2 (low)
|
512k
|
MBZUAI Institute of Foundation Models
|
9
*
|
--
|
--
|
--
|
--
|
| |
Kimi Linear 48B A3B Instruct
|
1M
|
Kimi
|
9
*
|
--
|
--
|
--
|
--
|
| |
Llama 3.1 405B
|
128k
|
Meta
|
9
*
|
$3.13
|
75
|
2.46
|
9.17
|
| |
LFM2.5-8B-A1B
|
32.8k
|
Liquid AI
|
8
*
|
$0.00
|
336
|
8.39
|
15.82
|
| |
Ring-flash-2.0
|
128k
|
InclusionAI
|
8
*
|
$0.18
|
--
|
--
|
--
|
| |
Olmo 3.1 32B Think
|
65.5k
|
Allen Institute for AI
|
8
*
|
$0.00
|
--
|
--
|
--
|
| |
Solar Pro 2
|
65.5k
|
Upstage
|
8
*
|
--
|
--
|
--
|
--
|
| |
Command A
|
256k
|
Cohere
|
8
*
|
$3.25
|
58
|
1.65
|
10.24
|
| |
Qwen3.5 2B
|
262k
|
Alibaba
|
8
|
--
|
--
|
--
|
--
|
| |
Llama 3.1 Nemotron 70B
|
128k
|
NVIDIA
|
8
*
|
$1.20
|
299
|
4.90
|
6.57
|
| |
NVIDIA Nemotron 3 Nano
|
1M
|
NVIDIA
|
7
*
|
$0.07
|
109
|
0.74
|
5.31
|
| |
NVIDIA Nemotron Nano 9B V2
|
131k
|
NVIDIA
|
7
*
|
$0.06
|
154
|
1.86
|
5.10
|
| |
Hermes 4 70B
|
128k
|
Nous Research
|
7
*
|
$0.16
|
95
|
1.36
|
6.65
|
| |
Ministral 3 3B
|
256k
|
Mistral
|
7
|
$0.10
|
171
|
0.64
|
3.56
|
| |
Granite 4.1 8B
|
131k
|
IBM
|
7
*
|
$0.06
|
117
|
0.79
|
5.04
|
| |
Sarvam 30B (high)
|
65.5k
|
Sarvam
|
7
*
|
$0.03
|
159
|
1.87
|
17.59
|
| |
Olmo 3.1 32B Instruct
|
65.5k
|
Allen Institute for AI
|
6
*
|
--
|
--
|
--
|
--
|
| |
Gemma 4 E2B
|
128k
|
Google
|
6
*
|
--
|
--
|
--
|
--
|
| |
R1 1776
|
128k
|
Perplexity
|
6
*
|
--
|
--
|
--
|
--
|
| |
Llama 3.2 90B (Vision)
|
128k
|
Meta
|
6
*
|
$1.38
|
59
|
1.16
|
9.66
|
| |
Phi-4 Mini
|
128k
|
Microsoft
|
6
|
$0.00
|
45
|
0.82
|
11.92
|
| |
EXAONE 4.0 32B
|
131k
|
LG AI Research
|
6
*
|
--
|
--
|
--
|
--
|
| |
Qwen3.5 2B
|
262k
|
Alibaba
|
6
|
--
|
--
|
--
|
--
|
| |
Qwen3.5 0.8B
|
262k
|
Alibaba
|
5
|
--
|
--
|
--
|
--
|
| |
DeepHermes 3 - Mistral 24B
|
32k
|
Nous Research
|
5
*
|
--
|
--
|
--
|
--
|
| |
Jamba 1.7 Large
|
256k
|
AI21 Labs
|
5
*
|
$2.60
|
58
|
1.44
|
10.01
|
| |
Granite 4.0 H Small
|
128k
|
IBM
|
5
*
|
$0.08
|
404
|
10.25
|
11.48
|
| |
Qwen3 Omni 30B A3B
|
65.5k
|
Alibaba
|
5
*
|
$0.32
|
91
|
2.01
|
7.48
|
| |
LFM2 24B A2B
|
32.8k
|
Liquid AI
|
5
*
|
--
|
--
|
--
|
--
|
| |
Phi-4
|
16k
|
Microsoft
|
5
*
|
$0.16
|
21
|
2.26
|
26.05
|
| |
Nova Micro
|
130k
|
Amazon
|
5
*
|
$0.03
|
290
|
1.08
|
2.81
|
| |
Granite 4.1 3B
|
131k
|
IBM
|
5
|
--
|
--
|
--
|
--
|
| |
NVIDIA Nemotron Nano 12B v2 VL
|
128k
|
NVIDIA
|
5
*
|
$0.24
|
216
|
1.11
|
3.43
|
| |
Phi-4 Multimodal
|
128k
|
Microsoft
|
5
*
|
$0.00
|
17
|
0.84
|
29.41
|
| |
Qwen3.5 0.8B
|
262k
|
Alibaba
|
4
*
|
--
|
--
|
--
|
--
|
| |
MiniCPM-V 4.6 1.3B
|
262k
|
OpenBMB
|
4
|
--
|
--
|
--
|
--
|
| |
Jamba Reasoning 3B
|
262k
|
AI21 Labs
|
4
*
|
--
|
--
|
--
|
--
|
| |
Reka Flash 3
|
128k
|
Reka AI
|
4
*
|
$0.26
|
--
|
--
|
--
|
| |
Olmo 3 7B Think
|
65.5k
|
Allen Institute for AI
|
4
*
|
--
|
--
|
--
|
--
|
| |
Molmo 7B-D
|
4.1k
|
Allen Institute for AI
|
4
*
|
--
|
--
|
--
|
--
|
| |
Ling-mini-2.0
|
131k
|
InclusionAI
|
4
*
|
--
|
--
|
--
|
--
|
| |
Llama 3.2 11B (Vision)
|
128k
|
Meta
|
3
*
|
$0.35
|
49
|
0.74
|
11.01
|
| |
Exaone 4.0 1.2B
|
64k
|
LG AI Research
|
3
*
|
--
|
--
|
--
|
--
|
| |
Olmo 3 7B
|
65.5k
|
Allen Institute for AI
|
3
*
|
$0.11
|
--
|
--
|
--
|
| |
Exaone 4.0 1.2B
|
64k
|
LG AI Research
|
3
*
|
--
|
--
|
--
|
--
|
| |
LFM2.5-1.2B-Thinking
|
32k
|
Liquid AI
|
3
*
|
--
|
--
|
--
|
--
|
| |
Jamba 1.7 Mini
|
258k
|
AI21 Labs
|
3
*
|
--
|
--
|
--
|
--
|
| |
LFM2 2.6B
|
32.8k
|
Liquid AI
|
3
*
|
--
|
--
|
--
|
--
|
| |
LFM2.5-1.2B-Instruct
|
32k
|
Liquid AI
|
3
*
|
--
|
--
|
--
|
--
|
| |
Granite 4.0 H 1B
|
128k
|
IBM
|
3
*
|
--
|
--
|
--
|
--
|
| |
Gemma 3 270M
|
32k
|
Google
|
2
*
|
--
|
--
|
--
|
--
|
| |
Apertus 70B Instruct
|
65.5k
|
Swiss AI Initiative
|
2
*
|
$1.03
|
--
|
--
|
--
|
| |
Granite 4.0 Micro
|
128k
|
IBM
|
2
*
|
--
|
--
|
--
|
--
|
| |
DeepHermes 3 - Llama-3.1 8B
|
128k
|
Nous Research
|
2
*
|
--
|
--
|
--
|
--
|
| |
Granite 4.0 1B
|
128k
|
IBM
|
2
*
|
--
|
--
|
--
|
--
|
| |
Molmo2-8B
|
36.9k
|
Allen Institute for AI
|
2
*
|
--
|
--
|
--
|
--
|
| |
LFM2 8B A1B
|
32.8k
|
Liquid AI
|
2
*
|
--
|
--
|
--
|
--
|
| |
LFM2.5-VL-1.6B
|
32k
|
Liquid AI
|
1
*
|
$0.00
|
424
|
9.37
|
10.55
|
| |
Granite 4.0 350M
|
32.8k
|
IBM
|
1
*
|
--
|
--
|
--
|
--
|
| |
Tiny Aya Global
|
8.19k
|
Cohere
|
1
*
|
$0.00
|
--
|
--
|
--
|
| |
Apertus 8B Instruct
|
65.5k
|
Swiss AI Initiative
|
1
*
|
$0.11
|
--
|
--
|
--
|
| |
Granite 4.0 H 350M
|
32.8k
|
IBM
|
1
*
|
--
|
--
|
--
|
--
|
| |
Claude Sonnet 5 (low)
|
1M
|
Anthropic
|
--
|
$1.54
|
60
|
1.43
|
9.81
|
| |
EXAONE 4.5 33B
|
262k
|
LG AI Research
|
--
|
--
|
--
|
--
|
--
|
| |
Claude Sonnet 5 (xhigh)
|
1M
|
Anthropic
|
--
|
$1.54
|
70
|
39.88
|
47.03
|
| |
Gemini 3 Deep Think
|
128k
|
Google
|
--
|
--
|
--
|
--
|
--
|
| |
Claude Sonnet 5 (high)
|
1M
|
Anthropic
|
--
|
$1.54
|
69
|
12.48
|
19.77
|
| |
Mi:dm K 2.5 Pro Preview
|
128k
|
Korea Telecom
|
--
|
--
|
--
|
--
|
--
|
| |
GPT-5.5 Pro (xhigh)
|
922k
|
OpenAI
|
--
|
--
|
--
|
--
|
--
|
| |
Cogito v2.1
|
128k
|
Deep Cogito
|
--
|
$1.25
|
90
|
0.83
|
28.57
|
| |
Claude Sonnet 5 (medium)
|
1M
|
Anthropic
|
--
|
$1.54
|
60
|
1.77
|
10.17
|
|
Maximum number of combined input & output tokens. Output tokens commonly have a significantly lower limit (varied by model).
Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).
Time to first token received, in seconds, after API request sent. For reasoning models which share reasoning tokens, this will be the first reasoning token. For models which do not support streaming, this represents time to receive the completion.
Price per token, shown in USD per million tokens. Price is a blend of cache hit, input, and output token prices using the selected ratio (default 7:2:1 cache-input-output).
Price per token generated by the model (received from the API), represented as USD per million Tokens.
Price per token included in the request/message sent to the API, represented as USD per million Tokens.
Metrics are 'live' and are based on the past 72 hours of measurements, measurements are taken 8 times a day for single requests and 2 times per day for parallel requests.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) currently ranks #1 on the Artificial Analysis LLM Leaderboard with an Intelligence Index score of 60, out of 142 models ranked.
The top models by Intelligence Index are: 1. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 2. GPT-5.6 Sol (max) (59), 3. GPT-5.6 Sol (xhigh) (58), 4. GPT-5.6 Sol (high) (56), 5. Claude Opus 4.8 (Adaptive Reasoning, Max Effort) (56).
Mercury 2 is the fastest at 762.8 tokens per second, followed by LFM2.5-VL-1.6B (424.2 t/s) and Granite 4.0 H Small (403.9 t/s).
Gemma 3n E4B Instruct is the most affordable at $0.02 per 1M tokens (blended 7:2:1 cache hit/input/output ratio), followed by Nova Micro ($0.03) and Sarvam 30B (high) ($0.03).
GLM-5.2 (max) is the highest-ranked open weights model with an Intelligence Index score of 51. There are 77 open weights models out of 142 total on the leaderboard.
The top open weights models by Intelligence Index are: 1. GLM-5.2 (max) (51), 2. MiniMax-M3 (44), 3. DeepSeek V4 Pro (Reasoning, Max Effort) (44).
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) leads among 100 reasoning models with an Intelligence Index score of 60. Reasoning models use extended thinking to solve complex problems before responding.
The leaderboard includes filters to narrow results by model type (reasoning vs non-reasoning), openness (open weights vs proprietary), and other criteria. You can also adjust prompt options to see how performance varies with different input lengths.
Click on any model name in the leaderboard to visit its dedicated comparison page with detailed charts covering intelligence, pricing, speed, latency, and more. You can also compare API providers for each model. View all models