AutomationBench-AA Cost Efficiency
Source-backed AutomationBench-AA scores plotted against estimated effective cost per 1M I/O tokens (a runtime estimate, not a published task cost). Explicit reasoning/nonreasoning siblings are connected where both are public. Color = provider.
AutomationBench-AA Cost Efficiency
Source-backed AutomationBench-AA scores plotted against estimated effective cost per 1M I/O tokens (a runtime estimate, not a published task cost). Explicit reasoning/nonreasoning siblings are connected where both are public. Color = provider.
This chart is part of the AutomationBench-AA benchmark page, which adds the model table, sources and how to read each view.
More / ExperimentalAutomationBench-AA vs Effective CostX = effective cost (log). Y = AutomationBench-AA. Exact AA v4.3.2 model/configuration; never substitute older benchmark versions.Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartIQ BenchmarksAutomationBench-AA ScoresExact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Knowledge Work using a provisional calibrated scale.Data: Artificial AnalysisOpen chartMore / ExperimentalSWE-Bench Verified Benchmark ScoresHistorical data. Vals archived cohort; retired from default Software Engineering domain. Each model's SWE-Bench Verified score. Color = provider.Data: Vals.ai, LLM StatsOpen chartMore / ExperimentalTerminal-Bench 2.0 Benchmark ScoresHistorical data. Superseded by 2.1 and 4.0. Legacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chartMore / ExperimentalTerminal-Bench 2.1 Benchmark ScoresHistorical data. Superseded in Programmatic Reasoning by Terminal-Bench 4.0. Each model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.Data: Artificial Analysis Terminal-Bench v2.1, Vals.ai Terminal-Bench 2.1, Artificial Analysis Intelligence Index v4.3.1Open chartMore / ExperimentalTerminal-Bench Hard Benchmark ScoresHistorical data. Historical hard subset; not a Terminal-Bench 4.0 result. Each model's Terminal-Bench Hard score. Color = provider.Data: Artificial AnalysisOpen chart