BullshitBench v2

A bird's-eye view of the benchmark: what it measures and every AI IQ chart built on it.

About BullshitBench v2

BullshitBench v2 has source-backed multi-effort configurations. The Cost Efficiency chart expands only those configurations; the score bar remains one canonical row per model.

BullshitBench v2 Cost Efficiency
Source-backed reasoning and nonreasoning configurations plotted against runtime effective cost. Each line is one canonical model. Color = provider.

How to read this chart

This chart compares source-backed BullshitBench v2 configurations with runtime effective cost. Multiple reasoning levels for the same model are connected; models with one available level remain standalone points. Up and to the left is better.

BullshitBench v2 Scores
Clear Pushback rate: share of attempts where a model clearly challenges a false premise instead of accepting nonsense. Color = provider.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
BullshitBench v2 vs Effective Cost
X = effective cost (log). Y = BullshitBench v2 Clear Pushback %. Color = provider.
Controls:
BullshitBench v21:1Cost

How to read this chart

Each point is a public model. The chart compares BullshitBench v2 % against Effective Cost (per 1M I/O Tokens), with color showing the model provider.