Flagship Lab Benchmark Correlation Heatmap
A pairwise benchmark correlation matrix using only the latest flagship model from each major US and Chinese lab.
Flagship Lab Benchmark Correlation Heatmap
Flagship Lab Benchmark Correlation Heatmap
A pairwise benchmark correlation matrix using only the latest flagship model from each major US and Chinese lab.
How to read this chart
Each cell compares two benchmark columns across models with both values present. Stronger color indicates a stronger Pearson correlation.
Data sources
ARC PrizeEpoch AI FrontierMathVals.aiArtificial AnalysisMathArenaManual source capturevals-model-pageArtificial Analysis model leaderboardVals.ai IOIArtificial Analysis Terminal-Bench v2.1Vals.ai Terminal-Bench 2.1Vals.ai ProgramBenchFrontierSWESWE-rebenchLLM StatsOSWorldLLM Stats OSWorld-VerifiedToolathlonLLM Stats ToolathlonScale LabsScale Labs MCP AtlasArena.ai Agent ArenaAgents' Last ExamEpoch AI SimpleQA VerifiedBullshitBench v2Scale Labs MultiChallengeKaggle FACTS Grounding leaderboardArena.aiEQ-Bench 3AttuneBench
More / ExperimentalTerminal-Bench 2.0 Benchmark ScoresLegacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chartMore / ExperimentalARC-AGI-3 Cost EfficiencyX = ARC Prize reported Cost (V3) (log). Y = ARC-AGI-3 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-2 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-2 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalARC-AGI-1 Cost EfficiencyX = ARC Prize reported cost/task (log). Y = ARC-AGI-1 %. Each line connects one model's published reasoning-effort levels, so the score-vs-cost tradeoff is visible per model. Defaults to the current model generation. Color = provider.Data: ARC PrizeOpen chartMore / ExperimentalMMLU-Pro Cost EfficiencyMMLU-Pro score versus runtime effective cost. Explicit reasoning and nonreasoning variants are connected; other published configurations remain standalone points. Color = provider.Data: Artificial Analysis, ARC Prize, Vals.ai +2 moreOpen chartMore / ExperimentalMMMU-Pro Cost EfficiencyX = Artificial Analysis reported cost per task (log). Y = MMMU-Pro score. Each line connects one model's published reasoning-effort levels from the shared configuration dataset; models with one available level remain standalone points. Color = provider.Data: Artificial Analysis, Artificial Analysis model leaderboardOpen chart