Mystery Game Puzzles Scores
Epoch’s 100 text-state puzzles in an undisclosed game. Best-next-move accuracy in a minimal agent scaffold; invalid or missing moves receive no credit. Prompts and transcripts are private. Contributes to Abstract Reasoning using a provisional calibrated scale.
Mystery Game Puzzles Scores
Epoch’s 100 text-state puzzles in an undisclosed game. Best-next-move accuracy in a minimal agent scaffold; invalid or missing moves receive no credit. Prompts and transcripts are private. Contributes to Abstract Reasoning using a provisional calibrated scale.
Mystery Game Puzzles Cost Efficiency
Source-backed Mystery Game Puzzles scores plotted against estimated effective cost per 1M I/O tokens (a runtime estimate, not a published task cost). Explicit reasoning/nonreasoning siblings are connected where both are public. Color = provider.
This chart is part of the Mystery Game Puzzles benchmark page, which adds the model table, sources and how to read each view.
IQ DimensionsAbstract Reasoning IQ Bell CurveEach model's estimated Abstract Reasoning IQ, placed on a standard normal IQ distribution.Data: ARC Prize, Epoch AI, Epoch AIOpen chartCoreIQ Bell CurveEach model's estimated IQ, placed on a standard normal IQ distribution.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartIQ BenchmarksChess Puzzles ScoresEpoch’s 100 novel chess positions, presented as text (FEN). Exact best-next-move accuracy; no vision input. Chess knowledge and planning both contribute. Contributes to Abstract Reasoning using a provisional calibrated scale.Data: Epoch AIOpen chartIQ BenchmarksARC-AGI-2 ScoresEach model's ARC-AGI-2 score. Color = provider.Data: ARC PrizeOpen chartIQ BenchmarksARC-AGI-1 Benchmark ScoresEach model's ARC-AGI-1 score. Color = provider.Data: ARC PrizeOpen chartIQ BenchmarksARC-AGI-3 Benchmark ScoresEach model's ARC-AGI-3 score. Color = provider.Data: ARC PrizeOpen chart