BullshitBench v2 Scores
Clear Pushback rate: share of attempts where a model clearly challenges a false premise instead of accepting nonsense. Color = provider.
BullshitBench v2 Scores
BullshitBench v2 Scores
Clear Pushback rate: share of attempts where a model clearly challenges a false premise instead of accepting nonsense. Color = provider.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
Data sources
BullshitBench v2 vs Effective Cost
BullshitBench v2 vs Effective Cost
X = effective cost (log). Y = BullshitBench v2 Clear Pushback %. Color = provider.
Controls:
How to read this chart
Each point is a public model. The chart compares BullshitBench v2 % against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
IQ DimensionsReliability IQEach model's Reliability IQ plotted on a standard normal IQ distributionData: Epoch AI SimpleQA Verified, Artificial Analysis, BullshitBench v2 +2 moreOpen chartCoreAI Models on the IQ Bell CurveEach model's estimated IQ plotted on a standard normal IQ distributionData: ARC Prize, Epoch AI FrontierMath, Vals.ai +24 moreOpen chartIQ BenchmarksAA Long Chain Reasoning ScoresArtificial Analysis Long Chain Reasoning scores for long-document extraction, synthesis, and reasoning. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksAA Omniscience ScoresArtificial Analysis Omniscience scores for factual reliability. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksIFBench Benchmark ScoresInstruction-following and constraint-adherence scores. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksSimpleQA Verified ScoresCorrect-answer rate on Epoch AI's verified factuality set. Color = provider.Data: Epoch AI SimpleQA VerifiedOpen chart