Terminal-Bench 2.1 Benchmark Scores

Each model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.

Terminal-Bench 2.1 Benchmark Scores
Each model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.
Controls:

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Terminal-Bench 2.1 vs Effective Cost
X = effective cost (log). Y = Terminal-Bench 2.1 %. Color = provider.
Controls:
Terminal-Bench1:1Cost

How to read this chart

Each point is a public model. The chart compares Terminal-Bench 2.1 % against Effective Cost (per 1M I/O Tokens), with color showing the model provider.