Chess Puzzles
A bird's-eye view of the benchmark: what it measures and every AI IQ chart built on it.
About Chess Puzzles
Chess Puzzles evaluates spatial reasoning and planning on 100 novel FEN positions generated by Epoch AI. Each puzzle has one best move, and scores are exact-match accuracy. The efficiency view compares published model runs and connects Claude Opus 4.6's thinking-budget variants, with inference costs estimated from public evaluation-log token usage and AI IQ's canonical model pricing.
How to read this chart
Each line is one model run at several reasoning-effort levels on Chess Puzzles. Every point shows the score and estimated run cost at that effort, with cheaper runs on the left. Up and to the left is better; a line that stays flat while moving right is buying little extra score with the extra spend. The current generation is shown by default; use the Generation filter to include previous and older models.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
How to read this chart
Each point is a public model. The chart compares Chess Puzzles accuracy against Effective Cost (per 1M I/O Tokens), with color showing the model provider.