The AI Intelligence Leaderboard
Estimating the intelligence of every major AI model
How AI IQ estimates model intelligence
- We archive source captures from public benchmark leaderboards and extract only source-backed values
- We map each benchmark score to an implied IQ using calibrated difficulty curves
- We group scored benchmarks into seven dimensions: abstract, mathematical, academic, programmatic, computer use, reliability, and knowledge work
- We conservatively fill missing benchmark and dimension estimates only inside the scoring pipeline
- Every derived IQ averages all seven scored dimensions, so missing coverage cannot make a model look better by omission
Ranked composite IQ
The same composite IQ estimates as the bell curve above, shown as a ranked leaderboard. Each bar is a model, colored by provider.
Composite IQ averages all seven scored dimensions — abstract, mathematical, academic, programmatic, computer use, reliability, and knowledge work — so missing coverage cannot make a model look better by omission.
Effective cost & iso-curves
Effective cost on the X-axis is sticker price for 1M I/O Tokens × measured or imputed usage multiplier. 1M I/O Tokens means 1M input tokens plus 1M output tokens, priced at published positive input/output rates.
Usage multipliers use source-backed token and task-cost data first, then same-lineage predecessors, closest measured peers, and finally a 1× fallback only for models with positive pricing. Zero/free-only rows are not plotted as $0 cost.
Iso-curves trace lines of equal preference for IQ versus cost. The slider weights quality vs cost: center is 1:1, drag toward Cost to make cost matter more, or toward IQ to make quality matter more. Models above and to the right of a curve are strictly better.
Completion time by task shape
Completion time is time to last token (TTLT). It combines time to first answer token with the time required to stream the requested output length.
Use the input and output sliders to compare short prompts, long-context prompts, brief answers, and long generations using the same underlying response-time model as the chart gallery.
Iso-curves trace equal preference for IQ versus speed. Models farther up and to the left are better: higher estimated IQ with lower total response time.
Tracking frontier progress
Each dot is a model with a known release date and a derived IQ estimate. Models are positioned left-to-right by release date, so the chart shows how the frontier changes over time rather than just where models rank today.
Provider-colored lines connect each lab's flagship frontier checkpoints. Codex, mini, nano, flash, coder, and smaller open-weight variants are omitted so the chart tracks each lab's main offering rather than every SKU.
This view is most useful for spotting whether a new release is actually ahead of its direct predecessor, or whether source coverage and conservative imputations are shaping the comparison.
How to read this chart
Models are placed on a standard IQ-style distribution using Abstract Reasoning IQ. Higher scores appear farther to the right.
Benchmarks in this dimension
How to read this chart
Each point is a public model. The chart compares Abstract Reasoning IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Abstract Reasoning IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Mathematical Reasoning IQ. Higher scores appear farther to the right.
Benchmarks in this dimension
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Mathematical Reasoning IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Mathematical Reasoning IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Mathematical Reasoning IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Academic Reasoning IQ. Higher scores appear farther to the right.
Benchmarks in this dimension
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Academic Reasoning IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Academic Reasoning IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Academic Reasoning IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Programmatic Reasoning IQ. Higher scores appear farther to the right.
Benchmarks in this dimension
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Programmatic Reasoning IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Programmatic Reasoning IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Programmatic Reasoning IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Computer Use IQ. Higher scores appear farther to the right.
Benchmarks in this dimension
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Computer Use IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Computer Use IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Computer Use IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Reliability IQ. Higher scores appear farther to the right.
Benchmarks in this dimension
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Reliability IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Reliability IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Reliability IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Knowledge Work IQ. Higher scores appear farther to the right.
Benchmarks in this dimension
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Knowledge Work IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Knowledge Work IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Knowledge Work IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Software Engineering IQ. Higher scores appear farther to the right.
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Software Engineering IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Software Engineering IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Software Engineering IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using WebDev & Design IQ. Higher scores appear farther to the right.
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated WebDev & Design IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares WebDev & Design IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher WebDev & Design IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Emotional Reasoning IQ. Higher scores appear farther to the right.
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Emotional Reasoning IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Emotional Reasoning IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Emotional Reasoning IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Biology IQ. Higher scores appear farther to the right.
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Biology IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Biology IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Biology IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Cybersecurity IQ. Higher scores appear farther to the right.
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Cybersecurity IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Cybersecurity IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Cybersecurity IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using AI/ML Engineering IQ. Higher scores appear farther to the right.
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated AI/ML Engineering IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares AI/ML Engineering IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher AI/ML Engineering IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Finance IQ. Higher scores appear farther to the right.
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Finance IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Finance IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Finance IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Writing IQ. Higher scores appear farther to the right.
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Writing IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Writing IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Writing IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Engineering Design IQ. Higher scores appear farther to the right.
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Engineering Design IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Engineering Design IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Engineering Design IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Legal IQ. Higher scores appear farther to the right.
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Legal IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Legal IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Legal IQ checkpoint for that provider.
How to read this chart
Models are placed on a standard IQ-style distribution using Healthcare IQ. Higher scores appear farther to the right.
How to read this chart
Each bar is one model, colored by provider and ranked by its estimated Healthcare IQ. Defaults to current-generation models; use Filters to include previous generations.
How to read this chart
Each point is a public model. The chart compares Healthcare IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
How to read this chart
Each point is a public model's benchmark score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.
How to read this chart
Each dot is a model with a known release date. Provider lines track public flagship releases over time and advance only when a newer model sets a higher Healthcare IQ checkpoint for that provider.
Get the weekly AI model intelligence newsletter
New model launches, benchmark shifts, cost-performance winners, and practical guidance on which models are actually worth using.
Read on SubstackScored benchmarks, 7 dimensions
Each benchmark score is mapped to an implied IQ through an expected-score ladder at IQ 70, 85, 100, 115, 130, 145, and 160. Benchmarks are grouped into seven equally weighted dimensions: Abstract Reasoning, Mathematical Reasoning, Academic Reasoning, Programmatic Reasoning, Computer Use, Reliability, and Knowledge Work.
Hard, low-gameability benchmarks can assign high IQ to modest raw scores. Easier, saturated, or contamination-sensitive benchmarks require very high scores before reaching the 145+ range.
FrontierMath Tier 4, FrontierMath Tier 1-3, ProofBench, MathArena, and AIME are included in the Mathematical Reasoning dimension and shown as standalone benchmark charts on the IQ page.
Models need direct or compatible ancestral measurement evidence in at least 2 of 7 dimensions to receive a derived IQ. The coverage count includes direct measurements only. One source benchmark is enough for a dimension estimate; additional coverage raises confidence. Six dimensions use Bayesian estimates with uncertain predecessor priors; Abstract Reasoning retains its direct-evidence method with ARC-based fallbacks. Raw benchmark results remain measured values.
Direct predecessor lineage is used first when it is explicit. Remaining missing dimensions use a matched lower-quartile cap based on models with similar capability across the other dimensions.
How dimensions relate to composite IQ
Composite IQ is the equal-weight average of seven dimension estimates: Abstract Reasoning, Mathematical Reasoning, Academic Reasoning, Programmatic Reasoning, Computer Use, Reliability, and Knowledge Work. Derived scores always use all seven, with missing dimensions conservatively filled before averaging.
R² measures how much of the variance in composite IQ is explained by each dimension. A higher R² means that dimension is a stronger predictor of overall model intelligence.