Terminal-Bench 2.1 Benchmark Scores

Historical data. Superseded in Programmatic Reasoning by Terminal-Bench 4.0. Each model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.

Terminal-Bench 2.1 Benchmark Scores
Terminal-Bench 2.1 Cost Efficiency

This chart is part of the Terminal-Bench 2.1 benchmark page, which adds the model table, sources and how to read each view.