SWE-rebench Cost Efficiency
Source-matched reasoning-effort score and task-cost configurations. Each line is one canonical model. Color = provider.
SWE-rebench Cost Efficiency
SWE-rebench Cost Efficiency
Source-matched reasoning-effort score and task-cost configurations. Each line is one canonical model. Color = provider.
How to read this chart
Each line uses the shared model-configuration dataset to compare source-matched reasoning-effort levels on SWE-rebench. Every point shows the score and cost per problem for that exact configuration. Canonical score bars and bell curves remain one point per model.
Data sources
More / ExperimentalSWE-rebench vs Cost/TaskX = SWE-rebench reported cost/problem (log). Y = SWE-rebench resolved rate. Color = provider.Data: SWE-rebenchOpen chartIQ BenchmarksSWE-rebench Benchmark ScoresEach model's SWE-rebench resolved rate. Color = provider.Data: SWE-rebenchOpen chartMore / ExperimentalTerminal-Bench 2.0 Benchmark ScoresLegacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chartMore / ExperimentalARC-AGI-3 Cost EfficiencyX = ARC Prize reported Cost (V3) (log). Y = ARC-AGI-3 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-2 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-2 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-1 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-1 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chart