Terminal-Bench 4.0

A bird's-eye view of the benchmark: what it measures and every AI IQ chart built on it.

About Terminal-Bench 4.0

Terminal-Bench 4.0 has source-backed multi-effort configurations. The Cost Efficiency chart expands only those configurations; the score bar remains one canonical row per model.

Terminal-Bench 4.0 Cost Efficiency
Source-matched reasoning-effort score and task-cost configurations. Each line is one canonical model. Color = provider.

How to read this chart

Each line uses the shared model-configuration dataset to compare source-matched reasoning-effort levels on Terminal-Bench 4.0. Every point shows the score and AA cost per task for that exact configuration. Canonical score bars and bell curves remain one point per model.

Terminal-Bench 4.0 Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Programmatic Reasoning using a provisional calibrated scale.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Sep 22, 2026
Terminal-Bench 4.0 vs Effective Cost
X = effective cost (log). Y = Terminal-Bench 4.0. Exact AA v4.3.2 model/configuration; never substitute older benchmark versions.
Controls:
Terminal-Bench 4.01:1Cost

How to read this chart

Each point is a public model. The chart compares Terminal-Bench 4.0 (%) against Effective Cost (per 1M I/O Tokens), with color showing the model provider.