FACTS Grounding Scores
Google FACTS Grounding scores for long-form responses fully grounded in provided documents. Color = provider.
FACTS Grounding Scores
Google FACTS Grounding scores for long-form responses fully grounded in provided documents. Color = provider.
FACTS Grounding Cost Efficiency
Source-backed FACTS Grounding scores plotted against estimated effective cost per 1M I/O tokens (a runtime estimate, not a published task cost). Explicit reasoning/nonreasoning siblings are connected where both are public. Color = provider.
This chart is part of the FACTS Grounding benchmark page, which adds the model table, sources and how to read each view.
IQ DimensionsReliability IQ Bell CurveEach model's estimated Reliability IQ, placed on a standard normal IQ distribution.Data: Epoch AI SimpleQA Verified, Artificial Analysis, BullshitBench v2 +2 moreOpen chartCoreIQ Bell CurveEach model's estimated IQ, placed on a standard normal IQ distribution.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartIQ BenchmarksAA Omniscience ScoresArtificial Analysis Omniscience scores for factual reliability. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksAA Long Context Reasoning v1.1 ScoresExact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Reliability using a provisional calibrated scale.Data: Artificial AnalysisOpen chartIQ BenchmarksIFBench Benchmark ScoresInstruction-following and constraint-adherence scores. Color = provider.Data: Artificial Analysis, Artificial Analysis Intelligence Index v4.3.2Open chartIQ BenchmarksBullshitBench v2 ScoresClear Pushback rate: share of attempts where a model clearly challenges a false premise instead of accepting nonsense. Color = provider.Data: BullshitBench v2Open chart