MirrorCode Scores
Epoch ML,+Private,2L: 30 tasks, three repeats, seven days and 10B tokens per attempt. Average rate of completely solved programs; distinct from the paper cohort. Contributes to Programmatic Reasoning using a provisional calibrated scale.
MirrorCode Scores
Epoch ML,+Private,2L: 30 tasks, three repeats, seven days and 10B tokens per attempt. Average rate of completely solved programs; distinct from the paper cohort. Contributes to Programmatic Reasoning using a provisional calibrated scale.
MirrorCode Cost Efficiency
Source-backed MirrorCode scores plotted against estimated effective cost per 1M I/O tokens (a runtime estimate, not a published task cost). Explicit reasoning/nonreasoning siblings are connected where both are public. Color = provider.
This chart is part of the MirrorCode benchmark page, which adds the model table, sources and how to read each view.
IQ DimensionsProgrammatic Reasoning IQ Bell CurveEach model's estimated Programmatic Reasoning IQ, placed on a standard normal IQ distribution.Data: Vals.ai, Vals.ai LiveCodeBench, Vals.ai IOI V2 +4 moreOpen chartCoreIQ Bell CurveEach model's estimated IQ, placed on a standard normal IQ distribution.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartIQ BenchmarksTerminal-Bench 4.0 ScoresExact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Programmatic Reasoning using a provisional calibrated scale.Data: Artificial AnalysisOpen chartIQ BenchmarksLiveCodeBench Benchmark ScoresEach model's LiveCodeBench score. Color = provider.Data: Vals.ai, Vals.ai LiveCodeBenchOpen chartIQ BenchmarksProgramBench ScoresAlmost Resolved score: share of ProgramBench tasks passing at least 95% of hidden tests. Color = provider.Data: Vals.ai ProgramBenchOpen chartIQ BenchmarksIOI V2 Scores2024–2026 olympiad tasks, no interactive grading feedback. Version 2 only; separately calibrated from archived V1. Contributes to Programmatic Reasoning using a provisional calibrated scale.Data: Vals.ai IOI V2Open chart