IFBench Benchmark Scores
Instruction-following and constraint-adherence scores. Color = provider.
IFBench Benchmark Scores
IFBench Benchmark Scores
Instruction-following and constraint-adherence scores. Color = provider.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
Data sources
IFBench vs Effective Cost
IFBench vs Effective Cost
X = effective cost (log). Y = IFBench %. Color = provider.
Controls:
How to read this chart
Each point is a public model. The chart compares IFBench % against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
Data sources
IQ DimensionsReliability IQEach model's Reliability IQ plotted on a standard normal IQ distributionData: Epoch AI SimpleQA Verified, Artificial Analysis, BullshitBench v2 +2 moreOpen chartCoreAI Models on the IQ Bell CurveEach model's estimated IQ plotted on a standard normal IQ distributionData: ARC Prize, Epoch AI FrontierMath, Vals.ai +24 moreOpen chartIQ BenchmarksAA Long Chain Reasoning ScoresArtificial Analysis Long Chain Reasoning scores for long-document extraction, synthesis, and reasoning. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksAA Omniscience ScoresArtificial Analysis Omniscience scores for factual reliability. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksBullshitBench v2 ScoresClear Pushback rate: share of attempts where a model clearly challenges a false premise instead of accepting nonsense. Color = provider.Data: BullshitBench v2Open chartIQ BenchmarksSimpleQA Verified ScoresCorrect-answer rate on Epoch AI's verified factuality set. Color = provider.Data: Epoch AI SimpleQA VerifiedOpen chart