updated Mon Sept 28

LLM Leaderboard

This LLM leaderboard displays the latest public benchmark performance for SOTA model versions released after April 2024. The data comes from model providers as well as independently run evaluations by Vellum or the open-source community. We feature results from non-saturated benchmarks, excluding outdated benchmarks (e.g. MMLU).

Top models per tasks

Best in Reasoning (GPQA Diamond)

100%95%91%86%81%
96.2
Claude Sonnet 5
96
GPT-6 Astra
95.4
Claude 3 Opus
94.6
GPT-5.6 Sol
94.3
Gemini 3.1 Pro
Best in Reasoning (GPQA Diamond)
ModelScore
Claude Sonnet 596.2%
GPT-6 Astra96%
Claude 3 Opus95.4%
GPT-5.6 Sol94.6%
Gemini 3.1 Pro94.3%

Best in Agentic Coding (SWE Bench)

100%94%88%82%77%
96.2
GPT-5.6 Sol
95.5
Claude Mythos 5
95
Claude Fable 5
93
GPT-5.6 Luna
88.6
Claude Opus 4.8
Best in Agentic Coding (SWE Bench)
ModelScore
GPT-5.6 Sol96.2%
Claude Mythos 595.5%
Claude Fable 595%
GPT-5.6 Luna93%
Claude Opus 4.888.6%
New

Best for Work Automations (AutoBench)

50%38%25%13%0%
48.8
GLM-5.3-Flash
41.4
GPT-6 Astra
40
Claude Opus 5.5
33.2
GPT-6 Sol
31.4
Claude Fable 5.1
Best for Work Automations (AutoBench)
ModelScore
GLM-5.3-Flash48.8%
GPT-6 Astra41.4%
Claude Opus 5.540%
GPT-6 Sol33.2%
Claude Fable 5.131.4%
New

Best in Computer Use (OSWorld)

85%81%76%72%68%
85
Claude Fable 5
83.4
Claude Opus 4.8
81.2
Claude Sonnet 5
78.7
GPT-5.5
78.5
Claude Sonnet 4.6
Best in Computer Use (OSWorld)
ModelScore
Claude Fable 585%
Claude Opus 4.883.4%
Claude Sonnet 581.2%
GPT-5.578.7%
Claude Sonnet 4.678.5%
New

Best in Browsing (BrowseComp)

95%90%86%81%77%
92.2
GPT-5.6 Sol
91.5
GPT-6 Astra
91.2
Kimi K3
90.8
Claude Opus 5
88
Claude Fable 5
Best in Browsing (BrowseComp)
ModelScore
GPT-5.6 Sol92.2%
GPT-6 Astra91.5%
Kimi K391.2%
Claude Opus 590.8%
Claude Fable 588%
New

Best in Terminal Use (Terminal-Bench 2.1)

90%87%83%80%77%
89.4
Gemini 3.8 Flash
88.8
GPT-5.6 Sol
88.3
Kimi K3
88
Claude Mythos 5
85.8
Gemini 3.7 Flash
Best in Terminal Use (Terminal-Bench 2.1)
ModelScore
Gemini 3.8 Flash89.4%
GPT-5.6 Sol88.8%
Kimi K388.3%
Claude Mythos 588%
Gemini 3.7 Flash85.8%

Fastest and most affordable models

Fastest Models (Tokens/sec)

1Llama 4 Scout2600 t/s
2Llama 3.1 405b969 t/s
3GLM 5.2347 t/s
4Kimi K2.6342.6 t/s
5Kimi K2.5337.7 t/s

Lowest Latency (TTFT)

1GPT-5.3 Codex0.003s
2Nova Micro0.3s
3Llama 4 Scout0.33s
4Gemini 2.0 Flash0.34s
5GPT-4o mini0.35s

Cheapest Models (per 1M tokens)

1Nova Micro$0.04 / $0.14
2Gemini 1.5 Flash$0.075 / $0.3
3Gemini 2.0 Flash$0.1 / $0.4
4GPT-4.1 nano$0.1 / $0.4
5GPT-6 Luna$0.1 / $0.5

Compare models

Side-by-side comparison of the latest models released in the last 9 months.

vs
Claude Opus 5.5Claude Mythos 5.1
Context size10000001000000
Cutoff dateJun 20262026-06
I/O cost$4 / $20$10 / $50
Max output-1000000
Latency--
Speed--
Best Overall (HLE)
Claude Opus 5.5
67.7
Claude Mythos 5.1
65
Best in Terminal Use (Terminal-Bench 2.1)
Claude Opus 5.5
66.4
Claude Mythos 5.1
60.9
Best in Agentic Coding (SWE-Bench)
Claude Opus 5.5-
Claude Mythos 5.1-
Best in Reasoning (GPQA Diamond)
Claude Opus 5.5-
Claude Mythos 5.1-

Compare Personal AI harnesses

Compare with
Vellum
Hermes
OpenClaw
Claude Cowork
Hermes
Open source
MIT
MIT
Apache 2.0
Proprietary
MIT
Time to set up
Easy
Moderate
Difficult
Easy
Moderate
Native channels
iOS, MacOS, Web, Voice, Email, Telegram, Slack, CLI
CLI / TUI
CLI, MacOS, Web
CLI, MacOS, Windows, Web
CLI / TUI
Memory
Managed memory
SQLite + markdown — you build the memory stack
Basic memory, context loss
Limited
SQLite + markdown — you build the memory stack
Security
Built-in security
DIY
DIY
No sandboxing
DIY
Hosting
Cloud or self-hosted
Self-hosted only
Self-hosted only
Anthropic cloud
Self-hosted only
Native integrations
Managed OAuth connections
No managed connectors
No managed connectors
MCP only
No managed connectors
Schedules
Cron + Heartbeat
Cron + Heartbeat
Cron + Heartbeat
Cron only
Cron + Heartbeat
Pricing
Free + API costs, Paid plans available
Free + DIY Hosting Costs + API costs
Free + DIY Hosting Costs + API costs
Paid plans available + API costs
Free + DIY Hosting Costs + API costs

Model Comparison

ModelContext sizeCutoff dateI/O costMax outputLatencySpeed
Claude Opus 5.51000000Jun 2026$4 / $20---
Claude Fable 5.110000002026-06$10 / $501000000--
Claude Mythos 5.110000002026-06$10 / $501000000--
Claude Sonnet 5.51000000Jun 2026$2 / $10---
GPT-6 Astra--$10 / $50---
GLM-5.3-Flash1,000,000-$0.15 / $0.5---
Gemini 3.8 Flash1048576-$0.75 / $3.75---
GPT-5.41050000-$2.5 / $15---
GPT-5.4 mini400000-$0.75 / $4.5---
GPT-5.4 nano400000-$0.2 / $1.25---
Gemini 3.8 Flash Cyber------
GPT-5.4 Pro--$30 / $180---
GPT-5.6 Cyber------
GPT-6 Luna1050000May 2026$0.1 / $0.5---
GPT-6 Sol1050000Apr 2026$2 / $10---

Context window, cost and speed comparison

Models Context Window Input Cost / 1M tokens Output Cost / 1M tokens Speed (tokens/second) Latency
Claude Opus 5.51000000$4$20n/an/a
Claude Fable 5.11000000$10$50n/an/a
Claude Mythos 5.11000000$10$50n/an/a
Claude Sonnet 5.51000000$2$10n/an/a
GPT-6 Astran/a$10$50n/an/a
GLM-5.3-Flash1,000,000$0.15$0.5n/an/a
Gemini 3.8 Flash1048576$0.75$3.75n/an/a
GPT-5.41050000$2.5$15n/an/a
GPT-5.4 mini400000$0.75$4.5n/an/a
GPT-5.4 nano400000$0.2$1.25n/an/a
Gemini 3.8 Flash Cybern/an/an/an/an/a
GPT-5.4 Pron/a$30$180n/an/a
GPT-5.6 Cybern/an/an/an/an/a
GPT-6 Luna1050000$0.1$0.5n/an/a
GPT-6 Sol1050000$2$10n/an/a

Benchmark glossary

Humanity's Last Exam
A crowd-sourced exam of extremely hard questions spanning every academic discipline. Designed to be the final exam before superhuman AI.
GPQA Diamond
Graduate-level science questions curated by domain experts. Tests advanced reasoning across physics, chemistry, and biology.
SWE-Bench Verified
Real GitHub issues from popular Python repos that the model must resolve end-to-end. Measures agentic software engineering ability.
AutoBench
Automation benchmark evaluating a model's ability to complete real-world work automation tasks using tools and multi-step workflows.
OSWorld-Verified
Real-world computer use tasks requiring GUI interaction in desktop environments. Measures end-to-end task completion on a real OS.
BrowseComp
Agentic web search benchmark testing a model's ability to browse and extract information from the web to answer complex questions.
Terminal-Bench 2.1
Terminal and tool use benchmark evaluating a model's ability to execute multi-step tasks in a terminal environment.

The Personal AI you were promised

Get Started →