EQ-Bench 3 Benchmark Scores
AI-judged EQ-Bench 3 Elo, a style-sensitive emotional reasoning signal. Color = provider.
EQ-Bench 3 Benchmark Scores
EQ-Bench 3 Benchmark Scores
AI-judged EQ-Bench 3 Elo, a style-sensitive emotional reasoning signal. Color = provider.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
Data sources
EQ-Bench 3 vs Effective Cost
EQ-Bench 3 vs Effective Cost
X = effective cost (log). Y = EQ-Bench 3 Elo. Color = provider.
Controls:
How to read this chart
Each point is a public model. The chart compares EQ-Bench 3 Elo against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
Data sources
CoreAI Models on the IQ Bell CurveEach model's estimated IQ plotted on a standard normal IQ distributionData: ARC Prize, Epoch AI FrontierMath, Vals.ai +24 moreOpen chartEmotional & InteractionEmotional Reasoning (EQ)Emotional Reasoning domain scores derived from EQ-Bench 3 and AttuneBench, excluded from Composite IQ.Data: EQ-Bench 3, AttuneBenchOpen chartEmotional & InteractionArena.ai Overall Benchmark ScoresHuman-preference Arena.ai Overall Elo scores. Color = provider.Data: Arena.aiOpen chartEmotional & InteractionAttuneBench Benchmark ScoresDefault-mode AttuneBench Composite scores from real human-AI conversations. Color = provider.Data: AttuneBenchOpen chartCoreIQ vs EQX = diagnostic EQ. Y = IQ. Color = model provider.Data: EQ-Bench 3, AttuneBench, ARC Prize +26 moreOpen chartCostEQ vs Effective CostDiagnostic EQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals.ai +2 moreOpen chart