Chess Puzzles

A bird's-eye view of the benchmark: what it measures and every AI IQ chart built on it.

About Chess Puzzles

Chess Puzzles evaluates spatial reasoning and planning on 100 novel FEN positions generated by Epoch AI. Each puzzle has one best move, and scores are exact-match accuracy. The efficiency view compares published model runs and connects Claude Opus 4.6's thinking-budget variants, with inference costs estimated from public evaluation-log token usage and AI IQ's canonical model pricing.

Chess Puzzles Cost Efficiency
X = estimated inference cost for Epoch AI's 100-puzzle run (log), derived from source-reported token usage and canonical model pricing. Y = Chess Puzzles exact-match accuracy. Single configurations appear as points; Claude Opus 4.6's published thinking-budget runs form a connected effort curve. Color = provider.

How to read this chart

Each line is one model run at several reasoning-effort levels on Chess Puzzles. Every point shows the score and estimated run cost at that effort, with cheaper runs on the left. Up and to the left is better; a line that stays flat while moving right is buying little extra score with the extra spend. The current generation is shown by default; use the Generation filter to include previous and older models.

Chess Puzzles Benchmark Scores
Exact-match accuracy on Epoch AI's novel chess puzzles. Color = provider.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Chess Puzzles vs Effective Cost
X = effective cost (log). Y = Chess Puzzles accuracy. Color = provider.
Controls:
Chess Puzzles1:1Cost

How to read this chart

Each point is a public model. The chart compares Chess Puzzles accuracy against Effective Cost (per 1M I/O Tokens), with color showing the model provider.