ProofBench 1.1 Scores
Native Lean 4.25.2 verification including native_decide. Version 1.1 only; not comparable to the unversioned historical cohort. Contributes to Mathematical Reasoning using a provisional calibrated scale.
ProofBench 1.1 Scores
ProofBench 1.1 Scores
Native Lean 4.25.2 verification including native_decide. Version 1.1 only; not comparable to the unversioned historical cohort. Contributes to Mathematical Reasoning using a provisional calibrated scale.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
ProofBench 1.1 vs Effective Cost
ProofBench 1.1 vs Effective Cost
X = effective cost (log). Y = ProofBench 1.1. Native Lean 4.25.2 verification including native_decide. Version 1.1 only; not comparable to the unversioned historical cohort.
Controls:
How to read this chart
Each point is a public model. The chart compares ProofBench 1.1 (%) against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
IQ DimensionsMathematical Reasoning IQEach model's Mathematical Reasoning IQ plotted on a standard normal IQ distributionData: Epoch AI FrontierMath, ProofBench 1.1, MathArena +2 moreOpen chartCoreIQ Bell CurveEach model's estimated IQ, placed on a standard normal IQ distribution.Data: ARC Prize, Epoch AI FrontierMath, ProofBench 1.1 +24 moreOpen chartIQ BenchmarksFrontierMath Tier 1-3 Benchmark ScoresEach model's FrontierMath Tier 1-3 score. Color = provider.Data: Epoch AI FrontierMathOpen chartIQ BenchmarksFrontierMath Tier 4 Benchmark ScoresEach model's FrontierMath Tier 4 score. Color = provider.Data: Epoch AI FrontierMathOpen chartIQ BenchmarksAIME Benchmark ScoresEach model's AIME score. Color = provider.Data: Vals.ai, Artificial AnalysisOpen chartIQ BenchmarksMathArena Benchmark ScoresEach model's MathArena expected performance. Color = provider.Data: MathArenaOpen chart