OSWorld 2.0 Cost Efficiency
Source-backed reasoning and nonreasoning configurations plotted against runtime effective cost. Each line is one canonical model. Color = provider.
OSWorld 2.0 Cost Efficiency
OSWorld 2.0 Cost Efficiency
Source-backed reasoning and nonreasoning configurations plotted against runtime effective cost. Each line is one canonical model. Color = provider.
How to read this chart
This chart compares source-backed OSWorld 2.0 configurations with runtime effective cost. Multiple reasoning levels for the same model are connected; models with one available level remain standalone points. Up and to the left is better.
More / ExperimentalOSWorld 2.0 vs Effective CostX = effective cost (log). Y = OSWorld 2.0. Official OSWorld 2.0 leaderboard only: 108 long-horizon desktop tasks, 500-step budget, full task set, binary completion accuracy, best reasoning/tool configuration per model. Partial-credit scores and provider-reported launch figures are excluded.Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +2 moreOpen chartIQ BenchmarksOSWorld 2.0 ScoresOfficial OSWorld 2.0 leaderboard only: 108 long-horizon desktop tasks, 500-step budget, full task set, binary completion accuracy, best reasoning/tool configuration per model. Partial-credit scores and provider-reported launch figures are excluded. Contributes to Computer Use using a provisional calibrated scale.Data: OSWorld 2.0 official leaderboardOpen chartMore / ExperimentalSWE-Bench Verified Benchmark ScoresHistorical data. Vals archived cohort; retired from default Software Engineering domain. Each model's SWE-Bench Verified score. Color = provider.Data: Vals.ai, LLM StatsOpen chartMore / ExperimentalTerminal-Bench 2.0 Benchmark ScoresHistorical data. Superseded by 2.1 and 4.0. Legacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chartMore / ExperimentalTerminal-Bench 2.1 Benchmark ScoresHistorical data. Superseded in Programmatic Reasoning by Terminal-Bench 4.0. Each model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.Data: Artificial Analysis Terminal-Bench v2.1, Vals.ai Terminal-Bench 2.1, Artificial Analysis Intelligence Index v4.3.1Open chartMore / ExperimentalTerminal-Bench Hard Benchmark ScoresHistorical data. Historical hard subset; not a Terminal-Bench 4.0 result. Each model's Terminal-Bench Hard score. Color = provider.Data: Artificial AnalysisOpen chart