The AI Intelligence Leaderboard

Estimating the intelligence of every major AI model

AI Models on the IQ Bell Curve
Each model's estimated IQ plotted on a standard normal IQ distribution

How AI IQ estimates model intelligence

  1. We archive source captures from public benchmark leaderboards and extract only source-backed values
  2. We map each benchmark score to an implied IQ using calibrated difficulty curves
  3. We group scored benchmarks into six dimensions: abstract, mathematical, academic, programmatic, computer use, and reliability
  4. We conservatively fill missing benchmark and dimension estimates only inside the scoring pipeline
  5. Every derived IQ averages all six scored dimensions, so missing coverage cannot make a model look better by omission
Estimated IQ's of AI Models
Each model's composite AI IQ estimate, ranked highest to lowest. Color = provider.

Ranked composite IQ

The same composite IQ estimates as the bell curve above, shown as a ranked leaderboard. Each bar is a model, colored by provider.

Composite IQ averages all seven scored dimensions — abstract, mathematical, scientific, frontend engineering, backend engineering, computer use, and reliability — so missing coverage cannot make a model look better by omission.

IQ vs Effective Cost
Each model's estimated IQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).
IQ 1:1 Cost

Effective cost & iso-curves

Effective cost on the X-axis is sticker price for 1M I/O Tokens × measured or imputed usage multiplier. 1M I/O Tokens means 1M input tokens plus 1M output tokens, priced at published positive input/output rates.

Usage multipliers use source-backed token and task-cost data first, then same-lineage predecessors, closest measured peers, and finally a 1× fallback only for models with positive pricing. Zero/free-only rows are not plotted as $0 cost.

Iso-curves trace lines of equal preference for IQ versus cost. The slider weights quality vs cost: center is 1:1, drag toward Cost to make cost matter more, or toward IQ to make quality matter more. Models above and to the right of a curve are strictly better.

IQ vs Completion Time
Completion time is time to last token (TTLT): time to first answer token plus output tokens divided by throughput (TPS).

Completion time by task shape

Completion time is time to last token (TTLT). It combines time to first answer token with the time required to stream the requested output length.

Use the input and output sliders to compare short prompts, long-context prompts, brief answers, and long generations using the same underlying response-time model as the chart gallery.

Iso-curves trace equal preference for IQ versus speed. Models farther up and to the left are better: higher estimated IQ with lower total response time.

Frontier IQ Over Time
X = release date. Y = estimated IQ. Provider step-lines connect each provider's flagship frontier checkpoints over time.

Tracking frontier progress

Each dot is a model with a known release date and a derived IQ estimate. Models are positioned left-to-right by release date, so the chart shows how the frontier changes over time rather than just where models rank today.

Provider-colored lines connect each lab's flagship frontier checkpoints. Codex, mini, nano, flash, coder, and smaller open-weight variants are omitted so the chart tracks each lab's main offering rather than every SKU.

This view is most useful for spotting whether a new release is actually ahead of its direct predecessor, or whether source coverage and conservative imputations are shaping the comparison.

Mathematical Reasoning IQ
Each model's Mathematical Reasoning IQ plotted on a standard normal IQ distribution

What it measures

Multi-step quantitative reasoning, from competition problems to research-level proofs.

Mathematical Reasoning vs Effective Cost
X = effective cost (log). Y = Mathematical Reasoning IQ. Color = provider.

Mathematical Reasoning IQ vs effective cost

Each model’s Mathematical Reasoning IQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier), on a log scale.

Iso-curves trace lines of equal preference for Mathematical Reasoning IQ versus cost. Use the slider to weight quality against cost — models above and to the right of a curve are strictly better.

Academic Reasoning IQ
Each model's Academic Reasoning IQ plotted on a standard normal IQ distribution

What it measures

Expert academic reasoning across science, the humanities, and multimodal university-level material.

Academic Reasoning vs Effective Cost
X = effective cost (log). Y = Academic Reasoning IQ. Color = provider.

Academic Reasoning IQ vs effective cost

Each model’s Academic Reasoning IQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier), on a log scale.

Iso-curves trace lines of equal preference for Academic Reasoning IQ versus cost. Use the slider to weight quality against cost — models above and to the right of a curve are strictly better.

Abstract Reasoning IQ
Each model's Abstract Reasoning IQ plotted on a standard normal IQ distribution

What it measures

Fluid problem-solving on novel puzzles a model cannot have memorized — abstracting patterns from just a few examples.

Abstract Reasoning vs Effective Cost
X = effective cost (log). Y = Abstract Reasoning IQ. Color = provider.

Abstract Reasoning IQ vs effective cost

Each model’s Abstract Reasoning IQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier), on a log scale.

Iso-curves trace lines of equal preference for Abstract Reasoning IQ versus cost. Use the slider to weight quality against cost — models above and to the right of a curve are strictly better.

Programmatic Reasoning IQ
Each model's Programmatic Reasoning IQ plotted on a standard normal IQ distribution

What it measures

Using code as a general problem-solving medium: algorithms, decomposition, execution, debugging, and task completion.

Programmatic Reasoning vs Effective Cost
X = effective cost (log). Y = Programmatic Reasoning IQ. Color = provider.

Programmatic Reasoning IQ vs effective cost

Each model’s Programmatic Reasoning IQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier), on a log scale.

Iso-curves trace lines of equal preference for Programmatic Reasoning IQ versus cost. Use the slider to weight quality against cost — models above and to the right of a curve are strictly better.

Computer Use IQ
Each model's Computer Use IQ plotted on a standard normal IQ distribution

What it measures

Agentic operation of real tools and environments — terminals, browsers, and desktop apps.

Computer Use vs Effective Cost
X = effective cost (log). Y = Computer Use IQ. Color = provider.

Computer Use IQ vs effective cost

Each model’s Computer Use IQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier), on a log scale.

Iso-curves trace lines of equal preference for Computer Use IQ versus cost. Use the slider to weight quality against cost — models above and to the right of a curve are strictly better.

Reliability IQ
Each model's Reliability IQ plotted on a standard normal IQ distribution

What it measures

Following instructions precisely, staying factual, challenging false premises, and grounding answers in long source documents.

Reliability vs Effective Cost
X = effective cost (log). Y = Reliability IQ. Color = provider.

Reliability IQ vs effective cost

Each model’s Reliability IQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier), on a log scale.

Iso-curves trace lines of equal preference for Reliability IQ versus cost. Use the slider to weight quality against cost — models above and to the right of a curve are strictly better.

IQ vs Completion Time vs Cost in 3D
3D scatter: X = completion time (time to last token, TTLT; log, faster to the right), Y = IQ, Z = effective cost (log). Color = provider. Drag to rotate.

Three tradeoffs at once

Most charts pit two qualities against each other. This view holds all three of the practical tradeoffs in one space: how smart a model is, how fast it answers, and what it costs to run.

IQ rises on the vertical axis, faster models sit to the right, and effective cost runs back into the depth axis on a log scale. The ideal model lives up, right, and toward the front — high intelligence, quick responses, and low cost. Drag to rotate and find where each provider clusters.

IQ Methodology