AI Model Benchmarks
Every major closed- and open-weight foundation model, version by version, on the standard public benchmarks, with what each costs. Pick a benchmark, toggle models on and off, and build the comparison you need.
101
37
12
Oct 10, 2026
Benchmark
ViewWeights
Labs
Models
| Claude Opus 5.5 | 87.6% |
|---|---|
| GPT-6 Astra | 87.3% |
| Claude Fable 5.1 | 85% |
| Gemini 3.8 Flash | 81.3% |
| Kimi K3 | 80.9% |
| Muse Spark 1.3 | 79% |
| Grok 4.7 | 73.4% |
| GLM-5.3 | 71.5% |
| Gemini 3.1 Pro | 70.8% |
| Qwen3.8 Max | 67.4% |
| DeepSeek V4 Pro | 54.7% |
For journalists and researchers
Use these charts in your story
Every chart and number on this page is free to cite, screenshot and republish, in print or online, as long as you credit Lounge.ai and link to lounge.ai/benchmarks. Scores carry their source link and run configuration; the data is refreshed every two weeks (last on Oct 10, 2026). Need a custom cut, a quote or the full dataset? Email hello@lounge.ai.
Credit line
Source: Lounge.ai AI Model Benchmarks (lounge.ai/benchmarks), accessed October 10, 2026. https://lounge.ai/benchmarks
Embed a chart (links back automatically)
<a href="https://lounge.ai/benchmarks"><img src="https://lounge.ai/api/benchmarks/chart?b=tb" alt="Terminal-Bench 2.1: top AI models, from the Lounge.ai AI Model Benchmarks" width="1200" height="675" style="max-width:100%;height:auto" /></a> <p>Source: <a href="https://lounge.ai/benchmarks">Lounge.ai AI Model Benchmarks</a></p>
Download charts (PNG, 1200 × 675)
| Terminal-Bench 2.1 | Top 10 · dark | light | Best per lab | Open-weight |
| SWE-bench Verified | Top 10 · dark | light | Best per lab | Open-weight |
| GPQA Diamond | Top 10 · dark | light | Best per lab | Open-weight |
| Humanity's Last Exam | Top 10 · dark | light | Best per lab | Open-weight |
| ARC-AGI-2 | Top 10 · dark | light | Best per lab | Open-weight |
| AIME 2026 | Top 10 · dark | light | Best per lab | Open-weight |
| τ²-bench (Telecom) | Top 10 · dark | light | Best per lab | Open-weight |
| MMLU-Pro | Top 10 · dark | light | Best per lab | Open-weight |
| Artificial Analysis Intelligence Index | Top 10 · dark | light | Best per lab | Open-weight |
| LMArena Text | Top 10 · dark | light | Best per lab | Open-weight |