BullshitBench v2 vs Completion Time
Each model's BullshitBench v2 score against its estimated completion time (time to last token) for the selected input and output lengths. Timing is the site's per-model estimate, not a published benchmark run time. Color = provider.
BullshitBench v2 vs Completion Time
Each model's BullshitBench v2 score against its estimated completion time (time to last token) for the selected input and output lengths. Timing is the site's per-model estimate, not a published benchmark run time. Color = provider.
Controls:
This chart is part of the BullshitBench v2 benchmark page, which adds the model table, sources and how to read each view.
More / ExperimentalBullshitBench v2 Cost EfficiencySource-backed BullshitBench v2 scores plotted against estimated effective cost per 1M I/O tokens (a runtime estimate, not a published task cost). Explicit reasoning/nonreasoning siblings are connected where both are public. Color = provider.Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +2 moreOpen chartIQ BenchmarksBullshitBench v2 ScoresClear Pushback rate: share of attempts where a model clearly challenges a false premise instead of accepting nonsense. Color = provider.Data: BullshitBench v2Open chartResponse TimeIQ vs Completion TimeEach model's estimated IQ against its completion time, which is time to last token (TTLT). Drag the input (1K–10K tokens) and output (100–10K tokens) sliders to see how it changes with the workload.Data: AI IQ methodology, Artificial AnalysisOpen chartMore / ExperimentalSWE-Bench Verified Benchmark ScoresHistorical data. Vals archived cohort; retired from default Software Engineering domain. Each model's SWE-Bench Verified score. Color = provider.Data: Vals.ai, LLM StatsOpen chartMore / ExperimentalTerminal-Bench 2.0 Benchmark ScoresHistorical data. Superseded by 2.1 and 4.0. Legacy Terminal-Bench 2.0 scores retained for historical comparison. Color = provider.Data: Terminal-BenchOpen chartMore / ExperimentalTerminal-Bench 2.1 Benchmark ScoresHistorical data. Superseded in Programmatic Reasoning by Terminal-Bench 4.0. Each model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.Data: Artificial Analysis Terminal-Bench v2.1, Vals.ai Terminal-Bench 2.1, Artificial Analysis Intelligence Index v4.3.1Open chart