GPQA Diamond Cost Efficiency
Source-matched reasoning-effort score and task-cost configurations. Each line is one canonical model. Color = provider.
GPQA Diamond Cost Efficiency
GPQA Diamond Cost Efficiency
Source-matched reasoning-effort score and task-cost configurations. Each line is one canonical model. Color = provider.
How to read this chart
Each line uses the shared model-configuration dataset to compare source-matched reasoning-effort levels on GPQA Diamond. Every point shows the score and AA cost per task for that exact configuration. Canonical score bars and bell curves remain one point per model.
Data sources
More / ExperimentalGPQA Diamond vs Effective CostX = effective cost (log). Y = GPQA Diamond %. Color = provider.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chartIQ BenchmarksGPQA Diamond Benchmark ScoresEach model's GPQA Diamond score. Color = provider.Data: Artificial AnalysisOpen chartMore / ExperimentalTerminal-Bench 2.0 Benchmark ScoresLegacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chartMore / ExperimentalARC-AGI-3 Cost EfficiencyX = ARC Prize reported Cost (V3) (log). Y = ARC-AGI-3 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-2 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-2 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-1 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-1 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chart