Programmatic Reasoning

A bird's-eye view of the dimension: the main IQ chart, the effective-cost frontier, and the benchmark charts that feed it.

Programmatic Reasoning IQ
Each model's Programmatic Reasoning IQ plotted on a standard normal IQ distribution

How to read this chart

Models are placed on a standard IQ-style distribution using Programmatic Reasoning IQ. Higher scores appear farther to the right.

Benchmarks in this dimension

Programmatic Reasoning vs Effective Cost
X = effective cost (log). Y = Programmatic Reasoning IQ. Color = provider.
Controls:
Programmatic1:1Cost

How to read this chart

Each point is a public model. The chart compares Programmatic Reasoning IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.

LiveCodeBench Benchmark Scores
Each model's LiveCodeBench score. Color = provider.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Terminal-Bench 2.1 Benchmark Scores
Each model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

IOI Benchmark Scores
Vals.ai accuracy across the 2024 and 2025 International Olympiad in Informatics tasks. Color = provider.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
ProgramBench Scores
Almost Resolved score: share of ProgramBench tasks passing at least 95% of hidden tests. Color = provider.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

SWE-rebench Benchmark Scores
Each model's SWE-rebench resolved rate. Color = provider.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
FrontierSWE Scores
FrontierSWE leaderboard tracking for software-engineering skill at the edge of human ability. Color = provider.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources