ProofBench 1.1 Cost Efficiency
Source-backed reasoning and nonreasoning configurations plotted against runtime effective cost. Each line is one canonical model. Color = provider.
ProofBench 1.1 Cost Efficiency
ProofBench 1.1 Cost Efficiency
Source-backed reasoning and nonreasoning configurations plotted against runtime effective cost. Each line is one canonical model. Color = provider.
How to read this chart
This chart compares source-backed ProofBench 1.1 configurations with runtime effective cost. Multiple reasoning levels for the same model are connected; models with one available level remain standalone points. Up and to the left is better.
More / ExperimentalProofBench 1.1 vs Effective CostX = effective cost (log). Y = ProofBench 1.1. Native Lean 4.25.2 verification including native_decide. Version 1.1 only; not comparable to the unversioned historical cohort.Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +2 moreOpen chartIQ BenchmarksProofBench 1.1 ScoresNative Lean 4.25.2 verification including native_decide. Version 1.1 only; not comparable to the unversioned historical cohort. Contributes to Mathematical Reasoning using a provisional calibrated scale.Data: ProofBench 1.1Open chartMore / ExperimentalSWE-Bench Verified Benchmark ScoresHistorical data. Vals archived cohort; retired from default Software Engineering domain. Each model's SWE-Bench Verified score. Color = provider.Data: Vals.ai, LLM StatsOpen chartMore / ExperimentalTerminal-Bench 2.0 Benchmark ScoresHistorical data. Superseded by 2.1 and 4.0. Legacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chartMore / ExperimentalTerminal-Bench 2.1 Benchmark ScoresHistorical data. Superseded in Programmatic Reasoning by Terminal-Bench 4.0. Each model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.Data: Artificial Analysis Terminal-Bench v2.1, Vals.ai Terminal-Bench 2.1, Artificial Analysis Intelligence Index v4.3.1Open chartMore / ExperimentalTerminal-Bench Hard Benchmark ScoresHistorical data. Historical hard subset; not a Terminal-Bench 4.0 result. Each model's Terminal-Bench Hard score. Color = provider.Data: Artificial AnalysisOpen chart