ProgramBench Cost Efficiency
ProgramBench Almost Resolved versus runtime effective cost. Explicit reasoning and nonreasoning variants are connected; other published configurations remain standalone points. Color = provider.
ProgramBench Cost Efficiency
ProgramBench Cost Efficiency
ProgramBench Almost Resolved versus runtime effective cost. Explicit reasoning and nonreasoning variants are connected; other published configurations remain standalone points. Color = provider.
How to read this chart
This chart compares source-backed ProgramBench configurations with runtime effective cost. Multiple reasoning levels for the same model are connected; models with one available level remain standalone points. Up and to the left is better.
More / ExperimentalProgramBench vs Effective CostX = effective cost (log). Y = ProgramBench Almost Resolved. Color = provider.Data: Artificial Analysis, ARC Prize, Vals.ai +1 moreOpen chartIQ BenchmarksProgramBench ScoresAlmost Resolved score: share of ProgramBench tasks passing at least 95% of hidden tests. Color = provider.Data: Vals.ai ProgramBenchOpen chartMore / ExperimentalTerminal-Bench 2.0 Benchmark ScoresLegacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chartMore / ExperimentalARC-AGI-3 Cost EfficiencyX = ARC Prize reported Cost (V3) (log). Y = ARC-AGI-3 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-2 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-2 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-1 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-1 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chart