ProofBench vs Effective Cost
Archived chart. Historical results combine benchmark versions and are excluded from IQ. Use the separately versioned ProofBench 1.1 for current comparisons.
ProofBench vs Effective Cost
ProofBench vs Effective Cost
Archived chart. Historical results combine benchmark versions and are excluded from IQ. Use the separately versioned ProofBench 1.1 for current comparisons.
Controls:
How to read this chart
Each point is a public model. The chart compares ProofBench against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
Archived chartProofBench Benchmark ScoresArchived chart. Historical results combine benchmark versions and are excluded from IQ. Use the separately versioned ProofBench 1.1 for current comparisons.Data: Vals.aiOpen chartArchived chartFrontierSWE ScoresArchived chart. V1 Dominance retired from Composite IQ and Software Engineering; V2 mean@5 requires a separate calibration.Data: FrontierSWEOpen chartMore / ExperimentalSWE-Bench Verified Benchmark ScoresHistorical data. Vals archived cohort; retired from default Software Engineering domain. Each model's SWE-Bench Verified score. Color = provider.Data: Vals.ai, LLM StatsOpen chartArchived chartIOI Benchmark ScoresArchived chart. IOI V1 archived; V2 uses a different harness and 2024–2026 tasks.Data: Vals.ai IOIOpen chartMore / ExperimentalTerminal-Bench 2.0 Benchmark ScoresHistorical data. Superseded by 2.1 and 4.0. Legacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chartMore / ExperimentalTerminal-Bench 2.1 Benchmark ScoresHistorical data. Superseded in Programmatic Reasoning by Terminal-Bench 4.0. Each model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.Data: Artificial Analysis Terminal-Bench v2.1, Vals.ai Terminal-Bench 2.1, Artificial Analysis Intelligence Index v4.3.1Open chart