AA Long Context Reasoning v1.1 Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Reliability using a provisional calibrated scale.
AA Long Context Reasoning v1.1 Scores
AA Long Context Reasoning v1.1 Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Reliability using a provisional calibrated scale.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
AA Long Context Reasoning v1.1 vs Effective Cost
AA Long Context Reasoning v1.1 vs Effective Cost
X = effective cost (log). Y = AA Long Context Reasoning v1.1. Exact AA v4.3.2 model/configuration; never substitute older benchmark versions.
Controls:
How to read this chart
Each point is a public model. The chart compares AA Long Context Reasoning v1.1 (%) against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
IQ DimensionsReliability IQEach model's Reliability IQ plotted on a standard normal IQ distributionData: Epoch AI SimpleQA Verified, Artificial Analysis, BullshitBench v2 +3 moreOpen chartCoreIQ Bell CurveEach model's estimated IQ, placed on a standard normal IQ distribution.Data: ARC Prize, Epoch AI FrontierMath, ProofBench 1.1 +24 moreOpen chartIQ BenchmarksAA Omniscience ScoresArtificial Analysis Omniscience scores for factual reliability. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksIFBench Benchmark ScoresInstruction-following and constraint-adherence scores. Color = provider.Data: Artificial Analysis, Artificial Analysis Intelligence Index v4.3.2Open chartIQ BenchmarksBullshitBench v2 ScoresClear Pushback rate: share of attempts where a model clearly challenges a false premise instead of accepting nonsense. Color = provider.Data: BullshitBench v2Open chartIQ BenchmarksSimpleQA Verified ScoresCorrect-answer rate on Epoch AI's verified factuality set. Color = provider.Data: Epoch AI SimpleQA VerifiedOpen chart