Chess Puzzles Cost Efficiency
X = estimated inference cost for Epoch AI's 100-puzzle run (log), derived from source-reported token usage and canonical model pricing. Y = Chess Puzzles exact-match accuracy. Single configurations appear as points; Claude Opus 4.6's published thinking-budget runs form a connected effort curve. Color = provider.
Chess Puzzles Cost Efficiency
Chess Puzzles Cost Efficiency
X = estimated inference cost for Epoch AI's 100-puzzle run (log), derived from source-reported token usage and canonical model pricing. Y = Chess Puzzles exact-match accuracy. Single configurations appear as points; Claude Opus 4.6's published thinking-budget runs form a connected effort curve. Color = provider.
How to read this chart
Each line is one model run at several reasoning-effort levels on Chess Puzzles. Every point shows the score and estimated run cost at that effort, with cheaper runs on the left. Up and to the left is better; a line that stays flat while moving right is buying little extra score with the extra spend. The current generation is shown by default; use the Generation filter to include previous and older models.
Data sources
More / ExperimentalChess Puzzles vs Effective CostX = effective cost (log). Y = Chess Puzzles accuracy. Color = provider.Data: Artificial Analysis, ARC Prize, Vals.ai +1 moreOpen chartIQ BenchmarksChess Puzzles Benchmark ScoresExact-match accuracy on Epoch AI's novel chess puzzles. Color = provider.Data: Epoch AI Chess PuzzlesOpen chartCostTask EfficiencyEach dot shows the inverse of the effective-cost usage multiplier. Higher means less price-adjusted task work: 2× is about half the median task effort. Source-backed multipliers are preferred; lineage, peer, and 1× fallbacks are labeled in tooltips.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chartCostIQ vs Effective CostEach model's estimated IQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: AI IQ methodologyOpen chartMore / ExperimentalTerminal-Bench 2.0 Benchmark ScoresLegacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chartMore / ExperimentalARC-AGI-3 Cost EfficiencyX = ARC Prize reported Cost (V3) (log). Y = ARC-AGI-3 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chart