updated Fri Sept 4
LLM Leaderboard
This LLM leaderboard displays the latest public benchmark performance for SOTA model versions released after April 2024. The data comes from model providers as well as independently run evaluations by Vellum or the open-source community. We feature results from non-saturated benchmarks, excluding outdated benchmarks (e.g. MMLU).
Best Overall (Humanity's Last Exam)
70%53%35%18%0%
| Model | Score |
|---|---|
| Claude Fable 5.1 | 65% |
| Claude Mythos 5.1 | 65% |
| Claude Opus 5 | 64.7% |
| Claude Mythos 5 | 64.5% |
| Claude Opus 4.8 | 57.9% |
| Claude Sonnet 5 | 57.4% |
| GPT-6 Astra | 57.2% |
| Kimi K3 | 56% |
| GLM-5.3-Flash | 55.3% |
| GLM 5.2 | 54.7% |
| Kimi K2.6 | 54% |
| DeepSeek V4 Flash | 51.6% |
| DeepSeek V4 Pro | 48.2% |
| GPT-5.6 Sol | 47.2% |
| Gemini 3 Pro | 45.8% |
| Kimi K2 Thinking | 44.9% |
| Gemini 3.1 Pro | 44.4% |
| GPT-5.5 Pro | 43.1% |
| GPT-5.5 | 41.4% |
| Gemini 3.5 Flash | 40.2% |
Top models per tasks
Best in Reasoning (GPQA Diamond)
100%95%91%86%81%
| Model | Score |
|---|---|
| Claude Sonnet 5 | 96.2% |
| GPT-6 Astra | 96% |
| Claude 3 Opus | 95.4% |
| GPT-5.6 Sol | 94.6% |
| Gemini 3.1 Pro | 94.3% |
Best in Agentic Coding (SWE Bench)
100%94%88%82%77%
| Model | Score |
|---|---|
| GPT-5.6 Sol | 96.2% |
| Claude Mythos 5 | 95.5% |
| Claude Fable 5 | 95% |
| GPT-5.6 Luna | 93% |
| Claude Opus 4.8 | 88.6% |
New
Best for Work Automations (AutoBench)
50%38%25%13%0%
| Model | Score |
|---|---|
| GLM-5.3-Flash | 48.8% |
| GPT-6 Astra | 41.4% |
| Claude Fable 5.1 | 31.4% |
| Claude Mythos 5.1 | 31.4% |
| Gemini 3.7 Flash | 30.4% |
New
Best in Computer Use (OSWorld)
85%81%76%72%68%
| Model | Score |
|---|---|
| Claude Fable 5 | 85% |
| Claude Opus 4.8 | 83.4% |
| Claude Sonnet 5 | 81.2% |
| GPT-5.5 | 78.7% |
| Claude Sonnet 4.6 | 78.5% |
New
Best in Browsing (BrowseComp)
95%90%86%81%77%
| Model | Score |
|---|---|
| GPT-5.6 Sol | 92.2% |
| GPT-6 Astra | 91.5% |
| Kimi K3 | 91.2% |
| Claude Opus 5 | 90.8% |
| Claude Fable 5 | 88% |
New
Best in Terminal Use (Terminal-Bench 2.1)
90%86%81%77%72%
| Model | Score |
|---|---|
| GPT-5.6 Sol | 88.8% |
| Kimi K3 | 88.3% |
| Claude Mythos 5 | 88% |
| Gemini 3.7 Flash | 85.8% |
| Claude Fable 5 | 84.3% |
Fastest and most affordable models
Fastest Models (Tokens/sec)
1
Llama 4 Scout2600 t/s
2
Llama 3.1 405b969 t/s
3
GLM 5.2347 t/s
GLM 5.2347 t/s4
Kimi K2.6342.6 t/s
Kimi K2.6342.6 t/s5
Kimi K2.5337.7 t/s
Kimi K2.5337.7 t/sLowest Latency (TTFT)
1
GPT-5.3 Codex0.003s
GPT-5.3 Codex0.003s2
Nova Micro0.3s
3
Llama 4 Scout0.33s
4
Gemini 2.0 Flash0.34s
5
GPT-4o mini0.35s
GPT-4o mini0.35sCheapest Models (per 1M tokens)
1
Nova Micro$0.04 / $0.14
2
Gemini 1.5 Flash$0.075 / $0.3
3
Gemini 2.0 Flash$0.1 / $0.4
4
GPT-4.1 nano$0.1 / $0.4
GPT-4.1 nano$0.1 / $0.45
Llama 4 Scout$0.11 / $0.34
Compare models
Side-by-side comparison of the latest models released in the last 9 months.
| Context size | 1000000 | 1000000 |
| Cutoff date | 2026-06 | 2026-06 |
| I/O cost | $10 / $50 | $10 / $50 |
| Max output | 1000000 | 1000000 |
| Latency | - | - |
| Speed | - | - |
Compare Personal AI harnesses
Model Comparison
| Model | Context size | Cutoff date | I/O cost | Max output | Latency | Speed |
|---|---|---|---|---|---|---|
| 1000000 | 2026-06 | $10 / $50 | 1000000 | - | - | |
| 1000000 | 2026-06 | $10 / $50 | 1000000 | - | - | |
| 1,000,000 | May 2026 | $5 / $25 | 128,000 | - | - | |
GPT-6 Astra | - | - | $10 / $50 | - | - | - |
Kimi K3 | 1,048,576 | - | $3 / $15 | - | 4.46s | 35.2 t/s |
GLM-5.3-Flash | 1,000,000 | - | $0.15 / $0.5 | - | - | - |
GLM 5.2 | 1,000,000 | Mar 2026 | $0.95 / $3 | 128,000 | 1.14s | 347 t/s |
| 1000000 | Jan 2026 | $0.14 / $0.28 | 384000 | 1.42s | 107.9 t/s | |
| 1,000,000 | Jan 2026 | $0.435 / $0.87 | 384,000 | 1.2s | 174.9 t/s | |
GPT-5.6 Sol | 1,050,000 | Feb 2026 | $5 / $30 | 128,000 | - | - |
| 1,000,000 | Jan 2026 | $2 / $12 | 65,536 | 20.34s | 136.2 t/s | |
| 1,000,000 | Jan 2026 | $1.5 / $9 | 65,536 | 23.16s | 175.4 t/s | |
| 1,048,576 | Mar 2026 | $0.75 / $3.75 | 65,536 | - | - | |
GPT-5.6 Luna | 1,050,000 | Feb 2026 | $0.2 / $1.2 | 128,000 | - | - |
GPT-5.6 Terra | 1,050,000 | Feb 2026 | $2 / $12 | 128,000 | - | - |
Context window, cost and speed comparison
| Models | Context Window | Input Cost / 1M tokens | Output Cost / 1M tokens | Speed (tokens/second) | Latency |
|---|---|---|---|---|---|
| 1000000 | $10 | $50 | n/a | n/a | |
| 1000000 | $10 | $50 | n/a | n/a | |
| 1,000,000 | $5 | $25 | n/a | n/a | |
GPT-6 Astra | n/a | $10 | $50 | n/a | n/a |
Kimi K3 | 1,048,576 | $3 | $15 | 35.2 t/s | 4.46 seconds |
GLM-5.3-Flash | 1,000,000 | $0.15 | $0.5 | n/a | n/a |
GLM 5.2 | 1,000,000 | $0.95 | $3 | 347 t/s | 1.14 seconds |
| 1000000 | $0.14 | $0.28 | 107.9 t/s | 1.42 seconds | |
| 1,000,000 | $0.435 | $0.87 | 174.9 t/s | 1.2 seconds | |
GPT-5.6 Sol | 1,050,000 | $5 | $30 | n/a | n/a |
| 1,000,000 | $2 | $12 | 136.2 t/s | 20.34 seconds | |
| 1,000,000 | $1.5 | $9 | 175.4 t/s | 23.16 seconds | |
| 1,048,576 | $0.75 | $3.75 | n/a | n/a | |
GPT-5.6 Luna | 1,050,000 | $0.2 | $1.2 | n/a | n/a |
GPT-5.6 Terra | 1,050,000 | $2 | $12 | n/a | n/a |
Benchmark glossary
- Humanity's Last Exam
- A crowd-sourced exam of extremely hard questions spanning every academic discipline. Designed to be the final exam before superhuman AI.
- GPQA Diamond
- Graduate-level science questions curated by domain experts. Tests advanced reasoning across physics, chemistry, and biology.
- SWE-Bench Verified
- Real GitHub issues from popular Python repos that the model must resolve end-to-end. Measures agentic software engineering ability.
- AutoBench
- Automation benchmark evaluating a model's ability to complete real-world work automation tasks using tools and multi-step workflows.
- OSWorld-Verified
- Real-world computer use tasks requiring GUI interaction in desktop environments. Measures end-to-end task completion on a real OS.
- BrowseComp
- Agentic web search benchmark testing a model's ability to browse and extract information from the web to answer complex questions.
- Terminal-Bench 2.1
- Terminal and tool use benchmark evaluating a model's ability to execute multi-step tasks in a terminal environment.