ARC-AGI-2
A bird's-eye view of the benchmark: what it measures and every AI IQ chart built on it.
About ARC-AGI-2
ARC-AGI-2 measures abstract reasoning: each task is a small grid puzzle that shows a few input-output examples of a hidden transformation rule, and the model must infer the rule and apply it to a new input. The puzzles are designed to be novel, so scores reflect on-the-spot reasoning rather than memorized knowledge. Every result is scored per task with a published cost per attempt, taken from the official ARC Prize leaderboard.
How to read this chart
Each line is one model run at several reasoning-effort levels on ARC-AGI-2. Every point shows the score and cost per task at that effort, with cheaper runs on the left. Up and to the left is better; a line that stays flat while moving right is buying little extra score with the extra spend. The current generation is shown by default; use the Generation filter to include previous and older models.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
How to read this chart
Each point is a public model. The chart compares ARC-AGI-2 % against Cost/Task, with color showing the model provider.