BrowseComp Benchmark Scores
Each model's BrowseComp score. Color = provider.
BrowseComp Benchmark Scores
BrowseComp Benchmark Scores
Each model's BrowseComp score. Color = provider.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
Data sources
BrowseComp vs Effective Cost
BrowseComp vs Effective Cost
X = effective cost (log). Y = BrowseComp %. Color = provider.
Controls:
How to read this chart
Each point is a public model. The chart compares BrowseComp % against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
Data sources
IQ DimensionsComputer Use IQEach model's Computer Use IQ plotted on a standard normal IQ distributionData: LLM Stats, OSWorld, LLM Stats OSWorld-Verified +6 moreOpen chartCoreAI Models on the IQ Bell CurveEach model's estimated IQ plotted on a standard normal IQ distributionData: ARC Prize, Epoch AI FrontierMath, Vals.ai +24 moreOpen chartIQ BenchmarksMCP Atlas Benchmark ScoresReal-world MCP tool-use pass rates. Color = provider.Data: Scale Labs, Scale Labs MCP AtlasOpen chartIQ BenchmarksArena.ai Agent Arena ScoresAggregate net improvement across real-world Agent Mode sessions. Color = provider.Data: Arena.ai Agent ArenaOpen chartIQ BenchmarksToolathlon Benchmark ScoresEach model's Toolathlon score. Color = provider.Data: LLM Stats, Toolathlon, LLM Stats ToolathlonOpen chartIQ BenchmarksAgents' Last Exam Benchmark ScoresFull / Overall pass rates from Agents' Last Exam. Color = provider.Data: Agents' Last ExamOpen chart