Knowledge Work

A bird's-eye view of the dimension: the main IQ chart, the effective-cost frontier, and the benchmark charts that feed it.

Knowledge Work IQ Bell Curve
Each model's estimated Knowledge Work IQ, placed on a standard normal IQ distribution.

How to read this chart

Models are placed on a standard IQ-style distribution using Knowledge Work IQ. Higher scores appear farther to the right.

Benchmarks in this dimension

Knowledge Work IQ vs Cost
Each model's estimated Knowledge Work IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).
Controls:
Knowledge Work1:1Cost

How to read this chart

Each point is a public model. The chart compares Knowledge Work IQ against Effective Cost (per 1M I/O Tokens), with color showing the model provider.

GDPval-AA v2.1 normalized score Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Knowledge Work using a provisional calibrated scale.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Oct 5, 2026
AA-Briefcase v1.1 Elo Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Knowledge Work using a provisional calibrated scale.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Oct 5, 2026
AutomationBench-AA Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Knowledge Work using a provisional calibrated scale.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Oct 2, 2026
GDP.pdf Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Knowledge Work using a provisional calibrated scale.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Oct 2, 2026
EnterpriseOps-Gym-AA Scores
Strict success on 1,117 oracle-mode enterprise workflows, graded against final database state using AA’s Stirrup harness. Contributes to Knowledge Work using a provisional calibrated scale.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Sep 29, 2026
AA-AnalystAgent Scores
Quantitative analysis from spreadsheets and documents. Headline pass^5: correct on every one of five attempts, not pass@5. Contributes to Knowledge Work using a provisional calibrated scale.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Sep 29, 2026