AI Model Benchmarks101

Every major closed- and open-weight foundation model, version by version, on the standard public benchmarks, with what each costs. Pick a benchmark, toggle models on and off, and build the comparison you need.
Model versions
101
28 current
Open-weight
37
64 closed
Labs
12
frontier + open labs
Last refreshed
Oct 10, 2026
updates every two weeks
Benchmark
View
Weights
Labs
Models 12
Terminal-Bench 2.1Agentic tasks in a live terminal (Vals AI run, Terminus-2 harness). Measures how well a model works a shell unattended. Lounge
0%25%50%75%100%Claude Opus 5.587.6%GPT-6 Astra87.3%Claude Fable 5.185%Gemini 3.8 Flash81.3%Kimi K380.9%Muse Spark 1.379%Grok 4.773.4%GLM-5.371.5%Gemini 3.1 Pro70.8%Qwen3.8 Max67.4%DeepSeek V4 Pro54.7%
No published Terminal-Bench score: GPT-6.1 Sol
lounge.ai/benchmarks
Claude Opus 5.587.6%
GPT-6 Astra87.3%
Claude Fable 5.185%
Gemini 3.8 Flash81.3%
Kimi K380.9%
Muse Spark 1.379%
Grok 4.773.4%
GLM-5.371.5%
Gemini 3.1 Pro70.8%
Qwen3.8 Max67.4%
DeepSeek V4 Pro54.7%
Model$ in / outTerminal-Bench▼SWE-benchGPQAHLEARC-AGI-2AIMEτ²-benchMMLU-ProAA IndexArena Elo
Claude Opus 5.5closed$4 / $2087.6%——67.7%91.7%—————
GPT-6 Astraclosed$10 / $5087.3%—96%57.2%95%—————
Claude Fable 5.1closed$10 / $5085%——60.9%90%—————
Claude Sonnet 5.5closed$2 / $1083.1%——56.9%——————
Gemini 3.8 Flashclosed$0.75 / $3.7581.3%———89.2%—————
Kimi K3open
2.8T / 104B active · Kimi K3 License
$3 / $1580.9%93.4%93.5%56%60.4%———57.11476
Muse Spark 1.3closed$1.25 / $4.2579%—————————
DeepSeek V4.1 Flashopen
552B / 16B active · MIT
$0.14 / $0.2874.5%—90.9%36.8%——————
Grok 4.7closed$2 / $673.4%—————————
GPT-6 Lunaclosed$0.1 / $0.573%———59.3%—————
GLM-5.3open
753B · GLM-5.3 License
$1.4 / $4.471.5%————————1475
Gemini 3.1 Proclosed$2 / $1270.8%—94.3%44.4%77.1%—95.6%—46.51480
Qwen3.8 Maxopen
2.4T / 95B active · Qwen3.8-Max License
$2 / $667.4%—92.6%43.6%—————1480
Kimi K2.7 Codeopen
1T / 32B active · Modified MIT
$0.6 / $2.567%—————90.1%—42—
GLM-5.3 Flashopen
320B / 18B active · MIT
—62.9%—————————
Qwen3.8 27Bopen
27B · Apache-2.0
—58.4%—89.2%30.8%——————
DeepSeek V4 Proopen
1.6T / 49B active · MIT
$0.435 / $0.8754.7%80.6%90.1%42.7%61.3%—96.2%87.5%44.3—
MiniMax M3open
428B / 23B active · MiniMax Community
$0.3 / $1.253.6%80.5%————88.9%—44.4—
Nemotron 3 Ultraopen
550B / 55B active · OpenMDW-1.1
—50.9%71.9%87%26.7%——83.3%86.8%37.8—
Mistral Medium 3.5open
128B · Modified MIT
—39%77.6%————94.2%—29.9—
Claude Haiku 5.5closed$0.1 / $0.5———45.9%——————
Claude Mythos 5.1closed$10 / $50——————————
GPT-6.1 Solclosed$2 / $10————94.2%—————
gpt-oss-120bopen
117B / 5.1B active · Apache-2.0
——62.4%————65.8%—23.8—
Gemini 4 ArgonclosedPreview$2 / $10——————————
Gemma 4 31Bopen
31B · Apache-2.0
——52%—26.5%——59.9%85.2%29.4—
Muse Glimmer 30Bopen
30B · Apache-2.0
——76%83.5%——94.7%————
Qwen3.8 Flash-Nextopen
125B / 6B active · Qwen Community 1.0
———91.7%35.9%——————
Mistral Large 4openPreview———————————
Mistral Small 4open
119B / 6.5B active · Apache-2.0
———————41.2%—19.6—
30 of 101 model versions shown · hover a score for its configuration, click it for the source · prices are USD per 1M tokens via each lab's API, standard tier
For journalists and researchers

Use these charts in your story

Every chart and number on this page is free to cite, screenshot and republish, in print or online, as long as you credit Lounge.ai and link to lounge.ai/benchmarks. Scores carry their source link and run configuration; the data is refreshed every two weeks (last on Oct 10, 2026). Need a custom cut, a quote or the full dataset? Email hello@lounge.ai.

Credit line
Source: Lounge.ai AI Model Benchmarks (lounge.ai/benchmarks), accessed October 10, 2026. https://lounge.ai/benchmarks
Embed a chart (links back automatically)
<a href="https://lounge.ai/benchmarks"><img src="https://lounge.ai/api/benchmarks/chart?b=tb" alt="Terminal-Bench 2.1: top AI models, from the Lounge.ai AI Model Benchmarks" width="1200" height="675" style="max-width:100%;height:auto" /></a>
<p>Source: <a href="https://lounge.ai/benchmarks">Lounge.ai AI Model Benchmarks</a></p>
Download charts (PNG, 1200 × 675)
Building a comparison above? “Download PNG” under the chart exports exactly the models you picked. Each image carries the Lounge mark and the source line.