MultiChallenge Scores
Multi-turn instruction retention, inference memory, editing, and self-coherence performance. Color = provider.
MultiChallenge Scores
MultiChallenge Scores
Multi-turn instruction retention, inference memory, editing, and self-coherence performance. Color = provider.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
Data sources
IQ DimensionsReliability IQEach model's Reliability IQ plotted on a standard normal IQ distributionData: Epoch AI SimpleQA Verified, Artificial Analysis, BullshitBench v2 +2 moreOpen chartCoreAI Models on the IQ Bell CurveEach model's estimated IQ plotted on a standard normal IQ distributionData: ARC Prize, Epoch AI Chess Puzzles, Epoch AI FrontierMath +25 moreOpen chartIQ BenchmarksAA Long Chain Reasoning ScoresArtificial Analysis Long Chain Reasoning scores for long-document extraction, synthesis, and reasoning. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksAA Omniscience ScoresArtificial Analysis Omniscience scores for factual reliability. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksIFBench Benchmark ScoresInstruction-following and constraint-adherence scores. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksBullshitBench v2 ScoresClear Pushback rate: share of attempts where a model clearly challenges a false premise instead of accepting nonsense. Color = provider.Data: BullshitBench v2Open chart