Terminal-Bench 2.1 Benchmark Scores
Each model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.
Terminal-Bench 2.1 Benchmark Scores
Terminal-Bench 2.1 Benchmark Scores
Each model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
Terminal-Bench 2.1 vs Effective Cost
Terminal-Bench 2.1 vs Effective Cost
X = effective cost (log). Y = Terminal-Bench 2.1 %. Color = provider.
Controls:
How to read this chart
Each point is a public model. The chart compares Terminal-Bench 2.1 % against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
IQ DimensionsProgrammatic Reasoning IQEach model's Programmatic Reasoning IQ plotted on a standard normal IQ distributionData: Vals.ai, Vals.ai IOI, Artificial Analysis Terminal-Bench v2.1 +4 moreOpen chartCoreAI Models on the IQ Bell CurveEach model's estimated IQ plotted on a standard normal IQ distributionData: ARC Prize, Epoch AI FrontierMath, Vals.ai +24 moreOpen chartIQ BenchmarksLiveCodeBench Benchmark ScoresEach model's LiveCodeBench score. Color = provider.Data: Vals.aiOpen chartIQ BenchmarksIOI Benchmark ScoresVals.ai accuracy across the 2024 and 2025 International Olympiad in Informatics tasks. Color = provider.Data: Vals.ai IOIOpen chartIQ BenchmarksProgramBench ScoresAlmost Resolved score: share of ProgramBench tasks passing at least 95% of hidden tests. Color = provider.Data: Vals.ai ProgramBenchOpen chartIQ BenchmarksSWE-rebench Benchmark ScoresEach model's SWE-rebench resolved rate. Color = provider.Data: SWE-rebenchOpen chart