Terminal-Bench 4.0 Scores

Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Programmatic Reasoning using a provisional calibrated scale.

Terminal-Bench 4.0 Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Programmatic Reasoning using a provisional calibrated scale.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Sep 22, 2026
Terminal-Bench 4.0 vs Effective Cost
X = effective cost (log). Y = Terminal-Bench 4.0. Exact AA v4.3.2 model/configuration; never substitute older benchmark versions.
Controls:
Terminal-Bench 4.01:1Cost

How to read this chart

Each point is a public model. The chart compares Terminal-Bench 4.0 (%) against Effective Cost (per 1M I/O Tokens), with color showing the model provider.