MMLU-Pro

A bird's-eye view of the benchmark: what it measures and every AI IQ chart built on it.

About MMLU-Pro

MMLU-Pro tests broad academic knowledge and reasoning with harder multiple-choice questions than the original MMLU. The efficiency view compares source-backed scores with effective cost and connects explicit reasoning/nonreasoning siblings where available.

MMLU-Pro Cost Efficiency
MMLU-Pro score versus runtime effective cost. Explicit reasoning and nonreasoning variants are connected; other published configurations remain standalone points. Color = provider.

How to read this chart

This chart compares source-backed MMLU-Pro configurations with runtime effective cost. Multiple reasoning levels for the same model are connected; models with one available level remain standalone points. Up and to the left is better.

Data sources
Artificial AnalysisARC PrizeVals.aiManual source capturevals-model-page
MMLU-Pro Benchmark Scores
Each model's MMLU-Pro score. Color = provider.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Manual source capturevals-model-page
MMLU-Pro vs Effective Cost
X = effective cost (log). Y = MMLU-Pro score. Color = provider.
Controls:
MMLU-Pro1:1Cost

How to read this chart

Each point is a public model. The chart compares MMLU-Pro score against Effective Cost (per 1M I/O Tokens), with color showing the model provider.

Data sources
Artificial AnalysisARC PrizeVals.aiManual source capturevals-model-page