SWE-rebench

A bird's-eye view of the benchmark: what it measures and every AI IQ chart built on it.

About SWE-rebench

SWE-rebench has source-backed multi-effort configurations. The Cost Efficiency chart expands only those configurations; the score bar remains one canonical row per model.

SWE-rebench Cost Efficiency
Source-matched reasoning-effort score and task-cost configurations. Each line is one canonical model. Color = provider.

How to read this chart

Each line uses the shared model-configuration dataset to compare source-matched reasoning-effort levels on SWE-rebench. Every point shows the score and cost per problem for that exact configuration. Canonical score bars and bell curves remain one point per model.

Data sources
SWE-rebench Benchmark Scores
Each model's SWE-rebench resolved rate. Color = provider.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
SWE-rebench vs Cost/Task
X = SWE-rebench reported cost/problem (log). Y = SWE-rebench resolved rate. Color = provider.
Controls:
SWE-rebench1:1Cost

How to read this chart

Each point is a public model. The chart compares SWE-rebench % against Cost/Task, with color showing the model provider.

Data sources