Chess Puzzles Cost Efficiency

X = estimated inference cost for Epoch AI's 100-puzzle run (log), derived from source-reported token usage and canonical model pricing. Y = Chess Puzzles exact-match accuracy. Single configurations appear as points; Claude Opus 4.6's published thinking-budget runs form a connected effort curve. Color = provider.

Chess Puzzles Cost Efficiency
X = estimated inference cost for Epoch AI's 100-puzzle run (log), derived from source-reported token usage and canonical model pricing. Y = Chess Puzzles exact-match accuracy. Single configurations appear as points; Claude Opus 4.6's published thinking-budget runs form a connected effort curve. Color = provider.

How to read this chart

Each line is one model run at several reasoning-effort levels on Chess Puzzles. Every point shows the score and estimated run cost at that effort, with cheaper runs on the left. Up and to the left is better; a line that stays flat while moving right is buying little extra score with the extra spend. The current generation is shown by default; use the Generation filter to include previous and older models.