updated Fri Sept 4

LLM Leaderboard

This LLM leaderboard displays the latest public benchmark performance for SOTA model versions released after April 2024. The data comes from model providers as well as independently run evaluations by Vellum or the open-source community. We feature results from non-saturated benchmarks, excluding outdated benchmarks (e.g. MMLU).

Top models per tasks

Best in Reasoning (GPQA Diamond)

100%95%91%86%81%
96.2
Claude Sonnet 5
96
GPT-6 Astra
95.4
Claude 3 Opus
94.6
GPT-5.6 Sol
94.3
Gemini 3.1 Pro
Best in Reasoning (GPQA Diamond)
ModelScore
Claude Sonnet 596.2%
GPT-6 Astra96%
Claude 3 Opus95.4%
GPT-5.6 Sol94.6%
Gemini 3.1 Pro94.3%

Best in Agentic Coding (SWE Bench)

100%94%88%82%77%
96.2
GPT-5.6 Sol
95.5
Claude Mythos 5
95
Claude Fable 5
93
GPT-5.6 Luna
88.6
Claude Opus 4.8
Best in Agentic Coding (SWE Bench)
ModelScore
GPT-5.6 Sol96.2%
Claude Mythos 595.5%
Claude Fable 595%
GPT-5.6 Luna93%
Claude Opus 4.888.6%
New

Best for Work Automations (AutoBench)

50%38%25%13%0%
48.8
GLM-5.3-Flash
41.4
GPT-6 Astra
31.4
Claude Fable 5.1
31.4
Claude Mythos 5.1
30.4
Gemini 3.7 Flash
Best for Work Automations (AutoBench)
ModelScore
GLM-5.3-Flash48.8%
GPT-6 Astra41.4%
Claude Fable 5.131.4%
Claude Mythos 5.131.4%
Gemini 3.7 Flash30.4%
New

Best in Computer Use (OSWorld)

85%81%76%72%68%
85
Claude Fable 5
83.4
Claude Opus 4.8
81.2
Claude Sonnet 5
78.7
GPT-5.5
78.5
Claude Sonnet 4.6
Best in Computer Use (OSWorld)
ModelScore
Claude Fable 585%
Claude Opus 4.883.4%
Claude Sonnet 581.2%
GPT-5.578.7%
Claude Sonnet 4.678.5%
New

Best in Browsing (BrowseComp)

95%90%86%81%77%
92.2
GPT-5.6 Sol
91.5
GPT-6 Astra
91.2
Kimi K3
90.8
Claude Opus 5
88
Claude Fable 5
Best in Browsing (BrowseComp)
ModelScore
GPT-5.6 Sol92.2%
GPT-6 Astra91.5%
Kimi K391.2%
Claude Opus 590.8%
Claude Fable 588%
New

Best in Terminal Use (Terminal-Bench 2.1)

90%86%81%77%72%
88.8
GPT-5.6 Sol
88.3
Kimi K3
88
Claude Mythos 5
85.8
Gemini 3.7 Flash
84.3
Claude Fable 5
Best in Terminal Use (Terminal-Bench 2.1)
ModelScore
GPT-5.6 Sol88.8%
Kimi K388.3%
Claude Mythos 588%
Gemini 3.7 Flash85.8%
Claude Fable 584.3%

Fastest and most affordable models

Fastest Models (Tokens/sec)

1Llama 4 Scout2600 t/s
2Llama 3.1 405b969 t/s
3GLM 5.2347 t/s
4Kimi K2.6342.6 t/s
5Kimi K2.5337.7 t/s

Lowest Latency (TTFT)

1GPT-5.3 Codex0.003s
2Nova Micro0.3s
3Llama 4 Scout0.33s
4Gemini 2.0 Flash0.34s
5GPT-4o mini0.35s

Cheapest Models (per 1M tokens)

1Nova Micro$0.04 / $0.14
2Gemini 1.5 Flash$0.075 / $0.3
3Gemini 2.0 Flash$0.1 / $0.4
4GPT-4.1 nano$0.1 / $0.4
5Llama 4 Scout$0.11 / $0.34

Compare models

Side-by-side comparison of the latest models released in the last 9 months.

vs
Claude Mythos 5.1Claude Fable 5.1
Context size10000001000000
Cutoff date2026-062026-06
I/O cost$10 / $50$10 / $50
Max output10000001000000
Latency--
Speed--
Best Overall (HLE)
Claude Mythos 5.1
65
Claude Fable 5.1
65
Best in Terminal Use (Terminal-Bench 2.1)
Claude Mythos 5.1
60.9
Claude Fable 5.1
55.8
Best in Agentic Coding (SWE-Bench)
Claude Mythos 5.1-
Claude Fable 5.1-
Best in Reasoning (GPQA Diamond)
Claude Mythos 5.1-
Claude Fable 5.1-

Compare Personal AI harnesses

Compare with
Vellum
Hermes
OpenClaw
Claude Cowork
Hermes
Open source
MIT
MIT
Apache 2.0
Proprietary
MIT
Time to set up
Easy
Moderate
Difficult
Easy
Moderate
Native channels
iOS, MacOS, Web, Voice, Email, Telegram, Slack, CLI
CLI / TUI
CLI, MacOS, Web
CLI, MacOS, Windows, Web
CLI / TUI
Memory
Managed memory
SQLite + markdown — you build the memory stack
Basic memory, context loss
Limited
SQLite + markdown — you build the memory stack
Security
Built-in security
DIY
DIY
No sandboxing
DIY
Hosting
Cloud or self-hosted
Self-hosted only
Self-hosted only
Anthropic cloud
Self-hosted only
Native integrations
Managed OAuth connections
No managed connectors
No managed connectors
MCP only
No managed connectors
Schedules
Cron + Heartbeat
Cron + Heartbeat
Cron + Heartbeat
Cron only
Cron + Heartbeat
Pricing
Free + API costs, Paid plans available
Free + DIY Hosting Costs + API costs
Free + DIY Hosting Costs + API costs
Paid plans available + API costs
Free + DIY Hosting Costs + API costs

Model Comparison

ModelContext sizeCutoff dateI/O costMax outputLatencySpeed
Claude Fable 5.110000002026-06$10 / $501000000--
Claude Mythos 5.110000002026-06$10 / $501000000--
Claude Opus 51,000,000May 2026$5 / $25128,000--
GPT-6 Astra--$10 / $50---
Kimi K31,048,576-$3 / $15-4.46s35.2 t/s
GLM-5.3-Flash1,000,000-$0.15 / $0.5---
GLM 5.21,000,000Mar 2026$0.95 / $3128,0001.14s347 t/s
DeepSeek V4 Flash1000000Jan 2026$0.14 / $0.283840001.42s107.9 t/s
DeepSeek V4 Pro1,000,000Jan 2026$0.435 / $0.87384,0001.2s174.9 t/s
GPT-5.6 Sol1,050,000Feb 2026$5 / $30128,000--
Gemini 3.1 Pro1,000,000Jan 2026$2 / $1265,53620.34s136.2 t/s
Gemini 3.5 Flash1,000,000Jan 2026$1.5 / $965,53623.16s175.4 t/s
Gemini 3.7 Flash1,048,576Mar 2026$0.75 / $3.7565,536--
GPT-5.6 Luna1,050,000Feb 2026$0.2 / $1.2128,000--
GPT-5.6 Terra1,050,000Feb 2026$2 / $12128,000--

Context window, cost and speed comparison

Models Context Window Input Cost / 1M tokens Output Cost / 1M tokens Speed (tokens/second) Latency
Claude Fable 5.11000000$10$50n/an/a
Claude Mythos 5.11000000$10$50n/an/a
Claude Opus 51,000,000$5$25n/an/a
GPT-6 Astran/a$10$50n/an/a
Kimi K31,048,576$3$1535.2 t/s4.46 seconds
GLM-5.3-Flash1,000,000$0.15$0.5n/an/a
GLM 5.21,000,000$0.95$3347 t/s1.14 seconds
DeepSeek V4 Flash1000000$0.14$0.28107.9 t/s1.42 seconds
DeepSeek V4 Pro1,000,000$0.435$0.87174.9 t/s1.2 seconds
GPT-5.6 Sol1,050,000$5$30n/an/a
Gemini 3.1 Pro1,000,000$2$12136.2 t/s20.34 seconds
Gemini 3.5 Flash1,000,000$1.5$9175.4 t/s23.16 seconds
Gemini 3.7 Flash1,048,576$0.75$3.75n/an/a
GPT-5.6 Luna1,050,000$0.2$1.2n/an/a
GPT-5.6 Terra1,050,000$2$12n/an/a

Benchmark glossary

Humanity's Last Exam
A crowd-sourced exam of extremely hard questions spanning every academic discipline. Designed to be the final exam before superhuman AI.
GPQA Diamond
Graduate-level science questions curated by domain experts. Tests advanced reasoning across physics, chemistry, and biology.
SWE-Bench Verified
Real GitHub issues from popular Python repos that the model must resolve end-to-end. Measures agentic software engineering ability.
AutoBench
Automation benchmark evaluating a model's ability to complete real-world work automation tasks using tools and multi-step workflows.
OSWorld-Verified
Real-world computer use tasks requiring GUI interaction in desktop environments. Measures end-to-end task completion on a real OS.
BrowseComp
Agentic web search benchmark testing a model's ability to browse and extract information from the web to answer complex questions.
Terminal-Bench 2.1
Terminal and tool use benchmark evaluating a model's ability to execute multi-step tasks in a terminal environment.

The Personal AI you were promised

GET STARTED