Agents' Last Exam

A bird's-eye view of the benchmark: what it measures and every AI IQ chart built on it.

About Agents' Last Exam

Agents' Last Exam has source-backed multi-effort configurations. The Cost Efficiency chart expands only those configurations; the score bar remains one canonical row per model.

Agents' Last Exam Cost Efficiency
Source-matched reasoning-effort score and task-cost configurations. Each line is one canonical model. Color = provider.

How to read this chart

Each line uses the shared model-configuration dataset to compare source-matched reasoning-effort levels on Agents' Last Exam. Every point shows the score and evaluation cost per task for that exact configuration. Canonical score bars and bell curves remain one point per model.

Data sources
Agents' Last Exam Benchmark Scores
Full / Overall pass rates from Agents' Last Exam. Color = provider.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Agents' Last Exam vs Effective Cost
X = effective cost (log). Y = Agents' Last Exam %. Color = provider.
Controls:
ALE1:1Cost

How to read this chart

Each point is a public model. The chart compares Agents' Last Exam % against Effective Cost (per 1M I/O Tokens), with color showing the model provider.