Charts
Canonical chart pages for sharing, deep-linking, and cross-navigation from IQ, EQ, Costs, and More
Core
CoreAI Models on the IQ Bell CurveEach model's estimated IQ plotted on a standard normal IQ distributionData: ARC Prize, Epoch AI FrontierMath, Vals.ai +24 moreOpen chartCoreEstimated IQ's of AI ModelsEach model's composite AI IQ estimate, ranked highest to lowest. Defaults to current-generation models. Color = provider.Data: ARC Prize, Epoch AI FrontierMath, Vals.ai +35 moreOpen chartCoreFrontier IQ Over TimeX = release date. Y = estimated IQ. Provider step-lines connect each provider's flagship frontier checkpoints over time.Data: ARC Prize, Epoch AI FrontierMath, Vals.ai +24 moreOpen chartCoreIQ vs EQX = diagnostic EQ. Y = IQ. Color = model provider.Data: EQ-Bench 3, AttuneBench, ARC Prize +26 moreOpen chartCoreIQ vs EQ vs Cost in 3D3D scatter: X = diagnostic EQ, Y = IQ, Z = effective cost (log). Color = provider. Drag to rotate.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chartCoreAI Models by Country Over TimeEach bar counts the AI models AI IQ tracks that had been released by that month, stacked by the country of the company that trained them.Data: AI IQ datasetOpen chartCoreAI Model Share by Country Over TimeThe national origin of tracked AI models as a share of the total — each month normalized to 100%, so the changing balance between countries is easy to read.Data: AI IQ datasetOpen chart
IQ Dimensions
IQ DimensionsAbstract Reasoning IQEach model's Abstract Reasoning IQ plotted on a standard normal IQ distributionData: ARC PrizeOpen chartIQ DimensionsMathematical Reasoning IQEach model's Mathematical Reasoning IQ plotted on a standard normal IQ distributionData: Epoch AI FrontierMath, Vals.ai, Artificial Analysis +1 moreOpen chartIQ DimensionsAcademic Reasoning IQEach model's Academic Reasoning IQ plotted on a standard normal IQ distributionData: Artificial Analysis, Manual source capture, vals-model-page +1 moreOpen chartIQ DimensionsProgrammatic Reasoning IQEach model's Programmatic Reasoning IQ plotted on a standard normal IQ distributionData: Vals.ai, Vals.ai IOI, Artificial Analysis Terminal-Bench v2.1 +4 moreOpen chartIQ DimensionsComputer Use IQEach model's Computer Use IQ plotted on a standard normal IQ distributionData: LLM Stats, OSWorld, LLM Stats OSWorld-Verified +6 moreOpen chartIQ DimensionsReliability IQEach model's Reliability IQ plotted on a standard normal IQ distributionData: Epoch AI SimpleQA Verified, Artificial Analysis, BullshitBench v2 +2 moreOpen chart
IQ Benchmarks
IQ BenchmarksARC-AGI-3 Benchmark ScoresEach model's ARC-AGI-3 score. Color = provider.Data: ARC PrizeOpen chartIQ BenchmarksARC-AGI-2 ScoresEach model's ARC-AGI-2 score. Color = provider.Data: ARC PrizeOpen chartIQ BenchmarksARC-AGI-1 Benchmark ScoresEach model's ARC-AGI-1 score. Color = provider.Data: ARC PrizeOpen chartIQ BenchmarksFrontierMath Tier 4 Benchmark ScoresEach model's FrontierMath Tier 4 score. Color = provider.Data: Epoch AI FrontierMathOpen chartIQ BenchmarksFrontierMath Tier 1-3 Benchmark ScoresEach model's FrontierMath Tier 1-3 score. Color = provider.Data: Epoch AI FrontierMathOpen chartIQ BenchmarksProofBench Benchmark ScoresEach model's ProofBench score. Color = provider.Data: Vals.aiOpen chartIQ BenchmarksMathArena Benchmark ScoresEach model's MathArena expected performance. Color = provider.Data: MathArenaOpen chartIQ BenchmarksAIME Benchmark ScoresEach model's AIME score. Color = provider.Data: Vals.ai, Artificial AnalysisOpen chartIQ BenchmarksArena.ai WebDev ScoresArena.ai WebDev Elo scores for front-end web development. Color = provider.Data: Arena.ai WebDevOpen chartIQ BenchmarksDesignArena Frontend ScoresDesignArena Frontend Elo for agentic front-end web apps. Color = provider.Data: DesignArena FrontendOpen chartIQ BenchmarksDesignArena Full Stack ScoresDesignArena Full Stack Elo for end-to-end frontend + backend builds. Color = provider.Data: DesignArena Full StackOpen chartIQ BenchmarksVibe Code Bench v1.1 ScoresVals.ai Vibe Code Bench v1.1 accuracy for building web applications from scratch. Color = provider.Data: Vals.ai Vibe Code Bench, Vals.ai Vibe Code Bench v1.1Open chartIQ BenchmarksSWE Marathon ScoresEach model's SWE Marathon aggregate Resolution Rate (Pass@1). Color = provider.Data: SWE MarathonOpen chartIQ BenchmarksAPEX-SWE ScoresMercor APEX-SWE overall Pass@1 scores for economically valuable software-engineering work. Color = provider.Data: Mercor APEX-SWEOpen chartIQ BenchmarksGSO-Bench ScoresEach model's GSO Opt@1 score for software optimization tasks. Color = provider.Data: GSO-BenchOpen chartIQ BenchmarksFrontierCode Diamond ScoresEach model's FrontierCode Diamond score for high-quality production-code tasks. Color = provider.Data: Cognition FrontierCodeOpen chartIQ BenchmarksFrontierSWE ScoresFrontierSWE leaderboard tracking for software-engineering skill at the edge of human ability. Color = provider.Data: FrontierSWEOpen chartIQ BenchmarksSWE-rebench Benchmark ScoresEach model's SWE-rebench resolved rate. Color = provider.Data: SWE-rebenchOpen chartIQ BenchmarksDeepSWE v1.1 Benchmark ScoresEach model's DeepSWE v1.1 Pass@1 score. Color = provider.Data: DeepSWE v1.1Open chartIQ BenchmarksKernelBench Hard Benchmark ScoresEach model's audited KernelBench Hard pass rate. Color = provider.Data: KernelBench HardOpen chartIQ BenchmarksKernelBench Mega Benchmark ScoresEach model's RTX PRO 6000 KernelBench Mega speedup over optimized PyTorch. Color = provider.Data: KernelBench MegaOpen chartIQ BenchmarksSWE-Bench Pro Benchmark ScoresEach model's SWE-Bench Pro percentage score. Color = provider.Data: Scale Labs, Scale Labs SWE-Bench Pro Public DatasetOpen chartIQ BenchmarksSWE-Bench Verified Benchmark ScoresEach model's SWE-Bench Verified score. Color = provider.Data: Vals.ai, LLM StatsOpen chartIQ BenchmarksLiveCodeBench Benchmark ScoresEach model's LiveCodeBench score. Color = provider.Data: Vals.aiOpen chartIQ BenchmarksIOI Benchmark ScoresVals.ai accuracy across the 2024 and 2025 International Olympiad in Informatics tasks. Color = provider.Data: Vals.ai IOIOpen chartIQ BenchmarksProgramBench ScoresAlmost Resolved score: share of ProgramBench tasks passing at least 95% of hidden tests. Color = provider.Data: Vals.ai ProgramBenchOpen chartIQ BenchmarksSciCode Benchmark ScoresEach model's SciCode score. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksCritPt Benchmark ScoresEach model's CritPt percentage score. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksHumanity's Last Exam Benchmark ScoresEach model's Humanity's Last Exam score. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksGPQA Diamond Benchmark ScoresEach model's GPQA Diamond score. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksMMLU-Pro Benchmark ScoresEach model's MMLU-Pro score. Color = provider.Data: Manual source capture, vals-model-pageOpen chartIQ BenchmarksMMMU-Pro Benchmark ScoresEach model's MMMU-Pro score. Color = provider.Data: Artificial Analysis, Artificial Analysis model leaderboardOpen chartIQ BenchmarksTerminal-Bench 2.1 Benchmark ScoresEach model's Terminal-Bench 2.1 pass@1 score, using Artificial Analysis as canonical and Vals.ai as fallback. Color = provider.Data: Artificial Analysis Terminal-Bench v2.1, Vals.ai Terminal-Bench 2.1Open chartIQ BenchmarksTerminal-Bench Hard Benchmark ScoresEach model's Terminal-Bench Hard score. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksBrowseComp Benchmark ScoresEach model's BrowseComp score. Color = provider.Data: LLM StatsOpen chartIQ BenchmarksOSWorld-Verified Benchmark ScoresEach model's OSWorld-Verified score. Color = provider.Data: LLM Stats, OSWorld, LLM Stats OSWorld-VerifiedOpen chartIQ BenchmarksToolathlon Benchmark ScoresEach model's Toolathlon score. Color = provider.Data: LLM Stats, Toolathlon, LLM Stats ToolathlonOpen chartIQ BenchmarksMCP Atlas Benchmark ScoresReal-world MCP tool-use pass rates. Color = provider.Data: Scale Labs, Scale Labs MCP AtlasOpen chartIQ BenchmarksAgents' Last Exam Benchmark ScoresFull / Overall pass rates from Agents' Last Exam. Color = provider.Data: Agents' Last ExamOpen chartIQ BenchmarksArena.ai Agent Arena ScoresAggregate net improvement across real-world Agent Mode sessions. Color = provider.Data: Arena.ai Agent ArenaOpen chartIQ BenchmarksIFBench Benchmark ScoresInstruction-following and constraint-adherence scores. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksAA Omniscience ScoresArtificial Analysis Omniscience scores for factual reliability. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksBullshitBench v2 ScoresClear Pushback rate: share of attempts where a model clearly challenges a false premise instead of accepting nonsense. Color = provider.Data: BullshitBench v2Open chartIQ BenchmarksSimpleQA Verified ScoresCorrect-answer rate on Epoch AI's verified factuality set. Color = provider.Data: Epoch AI SimpleQA VerifiedOpen chartIQ BenchmarksMultiChallenge ScoresMulti-turn instruction retention, inference memory, editing, and self-coherence performance. Color = provider.Data: Scale Labs MultiChallengeOpen chartIQ BenchmarksAA Long Chain Reasoning ScoresArtificial Analysis Long Chain Reasoning scores for long-document extraction, synthesis, and reasoning. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksFACTS Grounding ScoresGoogle FACTS Grounding scores for long-form responses fully grounded in provided documents. Color = provider.Data: Kaggle FACTS Grounding leaderboardOpen chart
Emotional & Interaction
Emotional & InteractionEmotional Reasoning (EQ)Emotional Reasoning domain scores derived from EQ-Bench 3 and AttuneBench, excluded from Composite IQ.Data: EQ-Bench 3, AttuneBenchOpen chartEmotional & InteractionArena.ai Overall Benchmark ScoresHuman-preference Arena.ai Overall Elo scores. Color = provider.Data: Arena.aiOpen chartEmotional & InteractionEQ-Bench 3 Benchmark ScoresAI-judged EQ-Bench 3 Elo, a style-sensitive emotional reasoning signal. Color = provider.Data: EQ-Bench 3Open chartEmotional & InteractionAttuneBench Benchmark ScoresDefault-mode AttuneBench Composite scores from real human-AI conversations. Color = provider.Data: AttuneBenchOpen chart
Cost
CostIQ vs Input/Output Token CostEach model's estimated IQ plotted against its published token price. Toggle between input and output price per 1M tokens.Data: AI IQ methodologyOpen chartCostIQ vs Blended Token CostEach model's estimated IQ plotted against the cost of a representative 1M-token blend. Toggle between a coding blend (cache-heavy — 800K cache-read + 100K input + 100K output) and a copywriting blend (output-heavy — 150K input + 850K output).Data: AI IQ methodologyOpen chartCostIQ vs Effective CostEach model's estimated IQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: AI IQ methodologyOpen chartCostEQ vs Effective CostDiagnostic EQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals.ai +2 moreOpen chartCostAI Models by CostPublished price for 1M I/O Tokens compared with effective cost after applying the measured or imputed task-usage multiplier.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chartCostInput Price vs Output PriceEach model positioned by its published token prices — input price (Y) against output price (X), both per 1M tokens on a log scale. The dashed line is the best fit across models, showing the typical output-to-input price relationship.Data: AI IQ datasetOpen chartCostTask EfficiencyEach dot shows the inverse of the effective-cost usage multiplier. Higher means less price-adjusted task work: 2× is about half the median task effort. Source-backed multipliers are preferred; lineage, peer, and 1× fallbacks are labeled in tooltips.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chartCostSticker Price vs Effective CostEach model's published sticker price for 1M I/O Tokens plotted against its effective cost after the measured or imputed task-usage multiplier. Points above the dashed line cost more per task than their sticker price suggests.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chart
Response Time
Response TimeIQ vs ThroughputThroughput is median output tokens per second (TPS). Each model's estimated IQ is plotted against how quickly it streams its answer after generation begins.Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeIQ vs LatencyLatency is time to first token (TTFT): how long a model takes to begin responding. Drag the input-length slider from 1K to 10K tokens to interpolate between AA's two measured TTFT anchors.Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeIQ vs Completion TimeCompletion time is time to last token (TTLT): time to first answer token plus output tokens ÷ throughput (TPS). Drag the input slider (1K–10K tokens) and output slider (100–10K tokens) to see how TTLT changes with the workload.Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeIQ vs Completion Time vs Cost in 3D3D scatter: X = completion time (time to last token, TTLT; log, faster to the right), Y = IQ, Z = effective cost (log). Color = provider. Drag to rotate.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chartResponse TimeAnimating Completion Time over IQ vs CostX = Effective Cost/1M I/O Tokens (log). Y = IQ. Clock spin rate represents completion time (time to last token, TTLT).Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeAnimating Cost over IQ vs Completion TimeX = completion time (time to last token, TTLT; log, reversed). Y = IQ. Effective Cost/1M I/O Tokens is shown as bills over a 10s period.Data: AI IQ methodology, Artificial AnalysisOpen chart