Charts
Canonical chart pages for sharing, deep-linking, and cross-navigation from IQ, EQ, Costs, and More
Core
CoreAI Models on the IQ Bell CurveEach model's estimated IQ plotted on a standard normal IQ distributionData: ARC Prize, Epoch AI FrontierMath, Vals.ai +27 moreOpen chartCoreFrontier IQ Over TimeX = release date. Y = estimated IQ. Provider step-lines connect each provider's flagship frontier checkpoints over time.Data: ARC Prize, Epoch AI FrontierMath, Vals.ai +27 moreOpen chartCoreIQ vs EQX = diagnostic EQ. Y = IQ. Color = model provider.Data: Arena.ai, EQ-Bench 3, AttuneBench +30 moreOpen chartCoreIQ vs EQ vs Cost in 3D3D scatter: X = diagnostic EQ, Y = IQ, Z = effective cost (log). Color = provider. Drag to rotate.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chartCoreAI Models by Country Over TimeEach bar counts the AI models AI IQ tracks that had been released by that month, stacked by the country of the company that trained them.Data: AI IQ datasetOpen chartCoreAI Model Share by Country Over TimeThe national origin of tracked AI models as a share of the total — each month normalized to 100%, so the changing balance between countries is easy to read.Data: AI IQ datasetOpen chart
IQ Dimensions
IQ DimensionsAbstract Reasoning IQEach model's Abstract Reasoning IQ plotted on a standard normal IQ distributionData: ARC PrizeOpen chartIQ DimensionsMathematical Reasoning IQEach model's Mathematical Reasoning IQ plotted on a standard normal IQ distributionData: Epoch AI FrontierMath, Vals.ai, Artificial Analysis +1 moreOpen chartIQ DimensionsScientific Reasoning IQEach model's Scientific Reasoning IQ plotted on a standard normal IQ distributionData: Artificial AnalysisOpen chartIQ DimensionsFrontend Engineering IQEach model's Frontend Engineering IQ plotted on a standard normal IQ distributionData: Arena.ai WebDev, DesignArena Frontend, DesignArena Full Stack +2 moreOpen chartIQ DimensionsBackend Engineering IQEach model's Backend Engineering IQ plotted on a standard normal IQ distributionData: Vals.ai, Cognition FrontierCode, Mercor APEX-SWE +7 moreOpen chartIQ DimensionsComputer Use IQEach model's Computer Use IQ plotted on a standard normal IQ distributionData: Artificial Analysis Terminal-Bench v2.1, Vals.ai Terminal-Bench 2.1, Artificial Analysis +8 moreOpen chartIQ DimensionsReliability IQEach model's Reliability IQ plotted on a standard normal IQ distributionData: Artificial Analysis, BullshitBench v2, Kaggle FACTS Grounding leaderboardOpen chart
IQ Benchmarks
IQ BenchmarksARC-AGI-3 Benchmark ScoresEach model's ARC-AGI-3 score. Color = provider.Data: ARC PrizeOpen chartIQ BenchmarksARC-AGI-2 ScoresEach model's ARC-AGI-2 score. Color = provider.Data: ARC PrizeOpen chartIQ BenchmarksARC-AGI-1 Benchmark ScoresEach model's ARC-AGI-1 score. Color = provider.Data: ARC PrizeOpen chartIQ BenchmarksFrontierMath Tier 4 Benchmark ScoresEach model's FrontierMath Tier 4 score. Color = provider.Data: Epoch AI FrontierMathOpen chartIQ BenchmarksFrontierMath Tier 1-3 Benchmark ScoresEach model's FrontierMath Tier 1-3 score. Color = provider.Data: Epoch AI FrontierMathOpen chartIQ BenchmarksProofBench Benchmark ScoresEach model's ProofBench score. Color = provider.Data: Vals.aiOpen chartIQ BenchmarksMathArena Benchmark ScoresEach model's MathArena expected performance. Color = provider.Data: MathArenaOpen chartIQ BenchmarksAIME Benchmark ScoresEach model's AIME score. Color = provider.Data: Vals.ai, Artificial AnalysisOpen chartIQ BenchmarksArena.ai WebDev ScoresArena.ai WebDev Elo scores for front-end web development. Color = provider.Data: Arena.ai WebDevOpen chartIQ BenchmarksDesignArena Frontend ScoresDesignArena Frontend Elo for agentic front-end web apps. Color = provider.Data: DesignArena FrontendOpen chartIQ BenchmarksDesignArena Full Stack ScoresDesignArena Full Stack Elo for end-to-end frontend + backend builds. Color = provider.Data: DesignArena Full StackOpen chartIQ BenchmarksVibe Code Bench ScoresVals.ai Vibe Code Bench v1.1 accuracy for building web applications from scratch. Color = provider.Data: Vals.ai Vibe Code Bench, Vals.ai Vibe Code Bench v1.1Open chartIQ BenchmarksSWE Marathon ScoresEach model's SWE Marathon aggregate Resolution Rate (Pass@1). Color = provider.Data: SWE MarathonOpen chartIQ BenchmarksAPEX-SWE ScoresMercor APEX-SWE overall Pass@1 scores for economically valuable software-engineering work. Color = provider.Data: Mercor APEX-SWEOpen chartIQ BenchmarksGSO-Bench ScoresEach model's GSO Opt@1 score for software optimization tasks. Color = provider.Data: GSO-BenchOpen chartIQ BenchmarksFrontierCode Diamond ScoresEach model's FrontierCode Diamond score for high-quality production-code tasks. Color = provider.Data: Cognition FrontierCodeOpen chartIQ BenchmarksFrontierSWE ScoresFrontierSWE leaderboard tracking for software-engineering skill at the edge of human ability. Color = provider.Data: FrontierSWEOpen chartIQ BenchmarksSWE-rebench Benchmark ScoresEach model's SWE-rebench resolved rate. Color = provider.Data: SWE-rebenchOpen chartIQ BenchmarksDeepSWE v1.1 Benchmark ScoresEach model's DeepSWE v1.1 Pass@1 score. Color = provider.Data: DeepSWE v1.1Open chartIQ BenchmarksDeepSWE Benchmark ScoresEach model's DeepSWE Pass@1 score. Color = provider.Data: DeepSWEOpen chartIQ BenchmarksKernelBench Hard Benchmark ScoresEach model's audited KernelBench Hard pass rate. Color = provider.Data: KernelBench HardOpen chartIQ BenchmarksKernelBench Mega Benchmark ScoresEach model's RTX PRO 6000 KernelBench Mega speedup over optimized PyTorch. Color = provider.Data: KernelBench MegaOpen chartIQ BenchmarksSWE-Bench Pro Benchmark ScoresEach model's SWE-Bench Pro percentage score. Color = provider.Data: Scale Labs, Scale Labs SWE-Bench Pro Public DatasetOpen chartIQ BenchmarksSWE-Bench Verified Benchmark ScoresEach model's SWE-Bench Verified score. Color = provider.Data: Vals.ai, LLM StatsOpen chartIQ BenchmarksLiveCodeBench Benchmark ScoresEach model's LiveCodeBench score. Color = provider.Data: Vals.aiOpen chartIQ BenchmarksSciCode Benchmark ScoresEach model's SciCode score. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksCritPt Benchmark ScoresEach model's CritPt percentage score. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksHumanity's Last Exam Benchmark ScoresEach model's Humanity's Last Exam score. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksGPQA Diamond Benchmark ScoresEach model's GPQA Diamond score. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksTerminal-Bench 2.1 Benchmark ScoresEach model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.Data: Artificial Analysis Terminal-Bench v2.1, Vals.ai Terminal-Bench 2.1Open chartIQ BenchmarksTerminal-Bench Hard Benchmark ScoresEach model's Terminal-Bench Hard score. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksBrowseComp Benchmark ScoresEach model's BrowseComp score. Color = provider.Data: LLM StatsOpen chartIQ BenchmarksOSWorld-Verified Benchmark ScoresEach model's OSWorld-Verified score. Color = provider.Data: LLM Stats, OSWorld, LLM Stats OSWorld-VerifiedOpen chartIQ BenchmarksToolathlon Benchmark ScoresEach model's Toolathlon score. Color = provider.Data: LLM Stats, Toolathlon, LLM Stats ToolathlonOpen chartIQ BenchmarksMCP Atlas Benchmark ScoresReal-world MCP tool-use pass rates. Color = provider.Data: Scale Labs, Scale Labs MCP AtlasOpen chartIQ BenchmarksAgents' Last Exam Benchmark ScoresFull / Overall pass rates from Agents' Last Exam. Color = provider.Data: Agents' Last ExamOpen chartIQ BenchmarksIFBench Benchmark ScoresInstruction-following and constraint-adherence scores. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksAA Omniscience ScoresArtificial Analysis Omniscience scores for factual reliability. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksBullshitBench v2 ScoresClear Pushback rate: share of attempts where a model clearly challenges a false premise instead of accepting nonsense. Color = provider.Data: BullshitBench v2Open chartIQ BenchmarksAA Long Chain Reasoning ScoresArtificial Analysis Long Chain Reasoning scores for long-document extraction, synthesis, and reasoning. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksFACTS Grounding ScoresGoogle FACTS Grounding scores for long-form responses fully grounded in provided documents. Color = provider.Data: Kaggle FACTS Grounding leaderboardOpen chart
EQ / Diagnostic
EQ / DiagnosticEmotional Reasoning (EQ)Diagnostic Emotional Reasoning scores, excluded from Composite IQ until the benchmark base is stronger.Data: EQ-Bench 3, Arena.ai, AttuneBenchOpen chartEQ / DiagnosticArena.ai Overall Benchmark ScoresHuman-preference Arena.ai Overall Elo scores. Color = provider.Data: Arena.aiOpen chartEQ / DiagnosticEQ-Bench 3 Benchmark ScoresAI-judged EQ-Bench 3 Elo, a style-sensitive emotional reasoning signal. Color = provider.Data: EQ-Bench 3Open chartEQ / DiagnosticAttuneBench Benchmark ScoresDefault-mode AttuneBench Composite scores from real human-AI conversations. Color = provider.Data: AttuneBenchOpen chart
Cost
CostIQ vs Input/Output Token CostEach model's estimated IQ plotted against its published token price. Toggle between input and output price per 1M tokens.Data: AI IQ methodologyOpen chartCostIQ vs Blended Token CostEach model's estimated IQ plotted against the cost of a representative 1M-token blend. Toggle between a coding blend (cache-heavy — 800K cache-read + 100K input + 100K output) and a copywriting blend (output-heavy — 150K input + 850K output).Data: AI IQ methodologyOpen chartCostIQ vs Effective CostEach model's estimated IQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: AI IQ methodologyOpen chartCostEQ vs Effective CostDiagnostic EQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals.ai +3 moreOpen chartCostAI Models by CostPublished price for 1M I/O Tokens compared with effective cost after applying the measured or imputed task-usage multiplier.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chartCostInput Price vs Output PriceEach model positioned by its published token prices — input price (Y) against output price (X), both per 1M tokens on a log scale. The dashed line is the best fit across models, showing the typical output-to-input price relationship.Data: AI IQ datasetOpen chartCostTask EfficiencyEach dot shows the inverse of the effective-cost usage multiplier. Higher means less price-adjusted task work: 2× is about half the median task effort. Source-backed multipliers are preferred; lineage, peer, and 1× fallbacks are labeled in tooltips.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chartCostSticker Price vs Effective CostEach model's published sticker price for 1M I/O Tokens plotted against its effective cost after the measured or imputed task-usage multiplier. Points above the dashed line cost more per task than their sticker price suggests.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chart
Response Time
Response TimeIQ vs Throughput (Tokens Per Second)Each model's estimated IQ plotted against its median output speed in tokens per secondData: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeIQ vs Latency (Time to First Token)Each model's estimated IQ plotted against its time to first token (TTFT). Drag the input-length slider from 1K to 10K tokens — dots glide between AA's two measured anchors (values in between are linearly interpolated).Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeIQ vs End-to-End Response TimeEach model's IQ against its end-to-end response time, built from its own components: time to first answer token (which grows with input length) plus output tokens ÷ tokens-per-second. Drag the input slider (1K–10K tokens) and output slider (100–10K tokens) to see how total latency scales.Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeIQ vs Speed vs Cost in 3D3D scatter: X = response time (log, faster to the right), Y = IQ, Z = effective cost (log). Color = provider. Drag to rotate.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chartResponse TimeAnimating Speed over IQ vs CostX = Effective Cost/1M I/O Tokens (log). Y = IQ. Spin rate = speed.Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeAnimating Cost over IQ vs SpeedX = response time (log, reversed). Y = IQ. Effective Cost/1M I/O Tokens shown as bills over a 10s period.Data: AI IQ methodology, Artificial AnalysisOpen chart