MathArena Cost Efficiency
Source-matched reasoning-effort score and task-cost configurations. Each line is one canonical model. Color = provider.
MathArena Cost Efficiency
MathArena Cost Efficiency
Source-matched reasoning-effort score and task-cost configurations. Each line is one canonical model. Color = provider.
How to read this chart
Each line uses the shared model-configuration dataset to compare source-matched reasoning-effort levels on MathArena. Every point shows the score and expected cost per problem for that exact configuration. Canonical score bars and bell curves remain one point per model.
Data sources
More / ExperimentalMathArena vs Effective CostX = effective cost (log). Y = MathArena. Color = provider.Data: Artificial Analysis, ARC Prize, Vals.ai +1 moreOpen chartIQ BenchmarksMathArena Benchmark ScoresEach model's MathArena expected performance. Color = provider.Data: MathArenaOpen chartMore / ExperimentalTerminal-Bench 2.0 Benchmark ScoresLegacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chartMore / ExperimentalARC-AGI-3 Cost EfficiencyX = ARC Prize reported Cost (V3) (log). Y = ARC-AGI-3 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-2 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-2 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-1 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-1 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chart