Every major closed- and open-weight foundation model, version by version, on the standard public benchmarks, with what each costs. Pick a benchmark, toggle models on and off, and build the comparison you need.
101 of 101 model versions shown · hover a score for its configuration, click it for the source · prices are USD per 1M tokens via each lab's API, standard tier
For journalists and researchers
Use these charts in your story
Every chart and number on this page is free to cite, screenshot and republish, in print or online, as long as you credit Lounge.ai and link to lounge.ai/benchmarks. Scores carry their source link and run configuration; the data is refreshed every two weeks (last on Oct 10, 2026). Need a custom cut, a quote or the full dataset? Email hello@lounge.ai.
Credit line
Source: Lounge.ai AI Model Benchmarks (lounge.ai/benchmarks), accessed October 10, 2026. https://lounge.ai/benchmarks
Embed a chart (links back automatically)
<a href="https://lounge.ai/benchmarks"><img src="https://lounge.ai/api/benchmarks/chart?b=tb" alt="Terminal-Bench 2.1: top AI models, from the Lounge.ai AI Model Benchmarks" width="1200" height="675" style="max-width:100%;height:auto" /></a>
<p>Source: <a href="https://lounge.ai/benchmarks">Lounge.ai AI Model Benchmarks</a></p>
Building a comparison above? “Download PNG” under the chart exports exactly the models you picked. Each image carries the Lounge mark and the source line.
How to read the scores
Most numbers are self-reported by the lab, read from the benchmark's own leaderboard; each cell links to its source and says whether a run used thinking, tools or a particular effort level. Gaps of a point or two are noise. Terminal-Bench is the Vals AI 2.1 run; HLE prefers no-tools figures; the AA Index is the July 2026 snapshot so rows stay comparable.
Benchmarks tell you where to start, not where to end up. Post what you measured on your own evals, and see which models other Lounge founders ship with on Tools.
Most numbers are self-reported by the lab, read from the benchmark's own leaderboard; each cell links to its source and says whether a run used thinking, tools or a particular effort level. Gaps of a point or two are noise. Terminal-Bench is the Vals AI 2.1 run; HLE prefers no-tools figures; the AA Index is the July 2026 snapshot so rows stay comparable.
Benchmarks tell you where to start, not where to end up. Post what you measured on your own evals, and see which models other Lounge founders ship with on Tools.