FrontierSWE Scores
Archived chart. V1 Dominance retired from Composite IQ and Software Engineering; V2 mean@5 requires a separate calibration.
FrontierSWE Scores
Archived chartFrontierSWE vs Effective CostArchived chart. V1 Dominance retired from Composite IQ and Software Engineering; V2 mean@5 requires a separate calibration.Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +2 moreOpen chartArchived chartFrontierSWE Cost EfficiencyArchived chart. V1 Dominance retired from Composite IQ and Software Engineering; V2 mean@5 requires a separate calibration.Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +2 moreOpen chartArchived chartProofBench Benchmark ScoresArchived chart. Historical results combine benchmark versions and are excluded from IQ. Use the separately versioned ProofBench 1.1 for current comparisons.Data: Vals.aiOpen chartMore / ExperimentalSWE-Bench Verified Benchmark ScoresHistorical data. Vals archived cohort; retired from default Software Engineering domain. Each model's SWE-Bench Verified score. Color = provider.Data: Vals.ai, LLM StatsOpen chartArchived chartIOI Benchmark ScoresArchived chart. IOI V1 archived; V2 uses a different harness and 2024–2026 tasks.Data: Vals.ai IOIOpen chartMore / ExperimentalTerminal-Bench 2.0 Benchmark ScoresHistorical data. Superseded by 2.1 and 4.0. Legacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chart