OSWorld 2.0

A bird's-eye view of the benchmark: what it measures and every AI IQ chart built on it.

About OSWorld 2.0

OSWorld 2.0 does not yet have two source-matched reasoning-effort score/cost rows for one model. The Cost Efficiency surface is reserved at the top of this page and will populate when qualifying data is imported.

OSWorld 2.0 Cost Efficiency
Source-backed reasoning and nonreasoning configurations plotted against runtime effective cost. Each line is one canonical model. Color = provider.

How to read this chart

This chart compares source-backed OSWorld 2.0 configurations with runtime effective cost. Multiple reasoning levels for the same model are connected; models with one available level remain standalone points. Up and to the left is better.

OSWorld 2.0 Scores
Official OSWorld 2.0 leaderboard only: 108 long-horizon desktop tasks, 500-step budget, full task set, binary completion accuracy, best reasoning/tool configuration per model. Partial-credit scores and provider-reported launch figures are excluded. Contributes to Computer Use using a provisional calibrated scale.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Sep 17, 2026
OSWorld 2.0 vs Effective Cost
X = effective cost (log). Y = OSWorld 2.0. Official OSWorld 2.0 leaderboard only: 108 long-horizon desktop tasks, 500-step budget, full task set, binary completion accuracy, best reasoning/tool configuration per model. Partial-credit scores and provider-reported launch figures are excluded.
Controls:
OSWorld 2.01:1Cost

How to read this chart

Each point is a public model. The chart compares OSWorld 2.0 (%) against Effective Cost (per 1M I/O Tokens), with color showing the model provider.