OSWorld-Verified Benchmark Scores
Each model's OSWorld-Verified score. Color = provider.
OSWorld-Verified Benchmark Scores
OSWorld-Verified Benchmark Scores
Each model's OSWorld-Verified score. Color = provider.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
Data sources
OSWorld-Verified vs Effective Cost
OSWorld-Verified vs Effective Cost
X = effective cost (log). Y = OSWorld-Verified %. Color = provider.
Controls:
How to read this chart
Each point is a public model. The chart compares OSWorld-Verified % against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
IQ DimensionsComputer Use IQEach model's Computer Use IQ plotted on a standard normal IQ distributionData: LLM Stats, OSWorld, LLM Stats OSWorld-Verified +6 moreOpen chartCoreAI Models on the IQ Bell CurveEach model's estimated IQ plotted on a standard normal IQ distributionData: ARC Prize, Epoch AI FrontierMath, Vals.ai +24 moreOpen chartIQ BenchmarksBrowseComp Benchmark ScoresEach model's BrowseComp score. Color = provider.Data: LLM StatsOpen chartIQ BenchmarksMCP Atlas Benchmark ScoresReal-world MCP tool-use pass rates. Color = provider.Data: Scale Labs, Scale Labs MCP AtlasOpen chartIQ BenchmarksArena.ai Agent Arena ScoresAggregate net improvement across real-world Agent Mode sessions. Color = provider.Data: Arena.ai Agent ArenaOpen chartIQ BenchmarksToolathlon Benchmark ScoresEach model's Toolathlon score. Color = provider.Data: LLM Stats, Toolathlon, LLM Stats ToolathlonOpen chart