Terminal-Bench 2.0 Benchmark Scores
Legacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.
Terminal-Bench 2.0 Benchmark Scores
Terminal-Bench 2.0 Benchmark Scores
Legacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
Data sources
Terminal-Bench 2.0 vs Effective Cost
Terminal-Bench 2.0 vs Effective Cost
X = effective cost (log). Y = Terminal-Bench 2.0 %. Color = provider.
Controls:
How to read this chart
Each point is a public model. The chart compares Terminal-Bench 2.0 % against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
More / ExperimentalARC-AGI-3 Cost EfficiencyX = ARC Prize reported Cost (V3) (log). Y = ARC-AGI-3 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-2 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-2 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-1 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-1 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalMMLU-Pro Cost EfficiencyMMLU-Pro score versus runtime effective cost. Explicit reasoning and nonreasoning variants are connected; other published configurations remain standalone points. Color = provider.Data: Artificial Analysis, ARC Prize, Vals.ai +2 moreOpen chartMore / ExperimentalMMMU-Pro Cost EfficiencyX = Artificial Analysis reported cost per task (log). Y = MMMU-Pro score. Each line connects one model's published reasoning-effort levels from the shared configuration dataset; models with one available level remain standalone points. Color = provider.Data: Artificial Analysis, Artificial Analysis model leaderboardOpen chartMore / ExperimentalIOI Cost EfficiencyIOI accuracy versus runtime effective cost. Explicit reasoning and nonreasoning variants are connected; other published configurations remain standalone points. Color = provider.Data: Artificial Analysis, ARC Prize, Vals.ai +1 moreOpen chart