Charts
Canonical chart pages for sharing, deep-linking, and cross-navigation from IQ, EQ, Costs, and More
Core
CoreIQ Bell CurveEach model's estimated IQ, placed on a standard normal IQ distribution.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartCoreIQ RankingsEach model's estimated IQ, ranked highest to lowest. Defaults to current-generation models.Data: ARC Prize, Epoch AI, Epoch AI +34 moreOpen chartCoreIQ Over TimeEach model's estimated IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartCoreIQ vs EQX = diagnostic EQ. Y = IQ. Color = model provider.Data: EQ-Bench 4 Elo, EQ-Bench 3, AttuneBench +29 moreOpen chartCoreIQ vs EQ vs Cost in 3D3D scatter: X = diagnostic EQ, Y = IQ, Z = effective cost (log). Color = provider. Drag to rotate.Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartCoreAI Models by Country Over TimeEach bar counts the AI models AI IQ tracks that had been released by that month, stacked by the country of their developer.Data: AI IQ datasetOpen chartCoreAI Model Share by Country Over TimeThe national origin of tracked AI models as a share of the total — each month normalized to 100%, so the changing balance between countries is easy to read.Data: AI IQ datasetOpen chart
IQ Dimensions
IQ DimensionsAbstract Reasoning IQ Bell CurveEach model's estimated Abstract Reasoning IQ, placed on a standard normal IQ distribution.Data: ARC Prize, Epoch AI, Epoch AIOpen chartIQ DimensionsMathematical Reasoning IQ Bell CurveEach model's estimated Mathematical Reasoning IQ, placed on a standard normal IQ distribution.Data: Epoch AI FrontierMath, ProofBench 1.1, MathArena +2 moreOpen chartIQ DimensionsAcademic Reasoning IQ Bell CurveEach model's estimated Academic Reasoning IQ, placed on a standard normal IQ distribution.Data: Artificial Analysis, Artificial Analysis Intelligence Index v4.3.2, Manual source capture +1 moreOpen chartIQ DimensionsProgrammatic Reasoning IQ Bell CurveEach model's estimated Programmatic Reasoning IQ, placed on a standard normal IQ distribution.Data: Vals.ai, Vals.ai LiveCodeBench, Vals.ai IOI V2 +4 moreOpen chartIQ DimensionsComputer Use IQ Bell CurveEach model's estimated Computer Use IQ, placed on a standard normal IQ distribution.Data: LLM Stats, LLM Stats BrowseComp, OSWorld +5 moreOpen chartIQ DimensionsReliability IQ Bell CurveEach model's estimated Reliability IQ, placed on a standard normal IQ distribution.Data: Epoch AI SimpleQA Verified, Artificial Analysis, BullshitBench v2 +2 moreOpen chartIQ DimensionsKnowledge Work IQ Bell CurveEach model's estimated Knowledge Work IQ, placed on a standard normal IQ distribution.Data: Artificial Analysis, Artificial Analysis, Artificial AnalysisOpen chartIQ DimensionsAbstract Reasoning IQ RankingsEach model's estimated Abstract Reasoning IQ, ranked highest to lowest. Defaults to current-generation models.Data: ARC Prize, Epoch AI, Epoch AIOpen chartIQ DimensionsAbstract Reasoning IQ vs SpeedEach model's estimated Abstract Reasoning IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: ARC Prize, Epoch AI, Epoch AI +1 moreOpen chartIQ DimensionsAbstract Reasoning IQ Over TimeEach model's estimated Abstract Reasoning IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartIQ DimensionsMathematical Reasoning IQ RankingsEach model's estimated Mathematical Reasoning IQ, ranked highest to lowest. Defaults to current-generation models.Data: Epoch AI FrontierMath, ProofBench 1.1, MathArena +2 moreOpen chartIQ DimensionsMathematical Reasoning IQ vs SpeedEach model's estimated Mathematical Reasoning IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Epoch AI FrontierMath, ProofBench 1.1, MathArena +2 moreOpen chartIQ DimensionsMathematical Reasoning IQ Over TimeEach model's estimated Mathematical Reasoning IQ by release date. Step-lines connect each provider's flagship models over time.Data: Epoch AI FrontierMath, ProofBench 1.1, MathArena +26 moreOpen chartIQ DimensionsAcademic Reasoning IQ RankingsEach model's estimated Academic Reasoning IQ, ranked highest to lowest. Defaults to current-generation models.Data: Artificial Analysis, Artificial Analysis Intelligence Index v4.3.2, Artificial Analysis +5 moreOpen chartIQ DimensionsAcademic Reasoning IQ vs SpeedEach model's estimated Academic Reasoning IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial Analysis, Artificial Analysis Intelligence Index v4.3.2, Manual source capture +1 moreOpen chartIQ DimensionsAcademic Reasoning IQ Over TimeEach model's estimated Academic Reasoning IQ by release date. Step-lines connect each provider's flagship models over time.Data: Artificial Analysis, Artificial Analysis Intelligence Index v4.3.2, Manual source capture +26 moreOpen chartIQ DimensionsProgrammatic Reasoning IQ RankingsEach model's estimated Programmatic Reasoning IQ, ranked highest to lowest. Defaults to current-generation models.Data: Vals.ai, Vals.ai LiveCodeBench, Vals.ai IOI V2 +4 moreOpen chartIQ DimensionsProgrammatic Reasoning IQ vs SpeedEach model's estimated Programmatic Reasoning IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Vals.ai, Vals.ai LiveCodeBench, Vals.ai IOI V2 +4 moreOpen chartIQ DimensionsProgrammatic Reasoning IQ Over TimeEach model's estimated Programmatic Reasoning IQ by release date. Step-lines connect each provider's flagship models over time.Data: Vals.ai, Vals.ai LiveCodeBench, Vals.ai IOI V2 +26 moreOpen chartIQ DimensionsComputer Use IQ RankingsEach model's estimated Computer Use IQ, ranked highest to lowest. Defaults to current-generation models.Data: LLM Stats, LLM Stats BrowseComp, LLM Stats +6 moreOpen chartIQ DimensionsComputer Use IQ vs SpeedEach model's estimated Computer Use IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: LLM Stats, LLM Stats BrowseComp, OSWorld +6 moreOpen chartIQ DimensionsComputer Use IQ Over TimeEach model's estimated Computer Use IQ by release date. Step-lines connect each provider's flagship models over time.Data: LLM Stats, LLM Stats BrowseComp, OSWorld +26 moreOpen chartIQ DimensionsReliability IQ RankingsEach model's estimated Reliability IQ, ranked highest to lowest. Defaults to current-generation models.Data: Epoch AI SimpleQA Verified, Artificial Analysis, BullshitBench v2 +3 moreOpen chartIQ DimensionsReliability IQ vs SpeedEach model's estimated Reliability IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Epoch AI SimpleQA Verified, Artificial Analysis, BullshitBench v2 +2 moreOpen chartIQ DimensionsReliability IQ Over TimeEach model's estimated Reliability IQ by release date. Step-lines connect each provider's flagship models over time.Data: Epoch AI SimpleQA Verified, Artificial Analysis, BullshitBench v2 +26 moreOpen chartIQ DimensionsKnowledge Work IQ RankingsEach model's estimated Knowledge Work IQ, ranked highest to lowest. Defaults to current-generation models.Data: Artificial Analysis, Artificial Analysis, Artificial AnalysisOpen chartIQ DimensionsKnowledge Work IQ vs SpeedEach model's estimated Knowledge Work IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial Analysis, Artificial Analysis, Artificial AnalysisOpen chartIQ DimensionsKnowledge Work IQ Over TimeEach model's estimated Knowledge Work IQ by release date. Step-lines connect each provider's flagship models over time.Data: Artificial Analysis, Artificial Analysis, Artificial Analysis +26 moreOpen chart
IQ Benchmarks
IQ BenchmarksARC-AGI-3 Benchmark ScoresEach model's ARC-AGI-3 score. Color = provider.Data: ARC PrizeOpen chartIQ BenchmarksARC-AGI-2 ScoresEach model's ARC-AGI-2 score. Color = provider.Data: ARC PrizeOpen chartIQ BenchmarksARC-AGI-1 Benchmark ScoresEach model's ARC-AGI-1 score. Color = provider.Data: ARC PrizeOpen chartIQ BenchmarksFrontierMath Tier 4 Benchmark ScoresEach model's FrontierMath Tier 4 score. Color = provider.Data: Epoch AI FrontierMathOpen chartIQ BenchmarksFrontierMath Tier 1-3 Benchmark ScoresEach model's FrontierMath Tier 1-3 score. Color = provider.Data: Epoch AI FrontierMathOpen chartIQ BenchmarksMathArena Benchmark ScoresEach model's MathArena expected performance. Color = provider.Data: MathArenaOpen chartIQ BenchmarksAIME Benchmark ScoresEach model's AIME score. Color = provider.Data: Vals.ai, Artificial AnalysisOpen chartIQ BenchmarksArena.ai WebDev ScoresArena.ai WebDev Elo scores for front-end web development. Color = provider.Data: Arena.ai WebDevOpen chartIQ BenchmarksDesignArena Frontend ScoresDesignArena Frontend Elo for agentic front-end web apps. Color = provider.Data: DesignArena FrontendOpen chartIQ BenchmarksDesignArena Full Stack ScoresDesignArena Full Stack Elo for end-to-end frontend + backend builds. Color = provider.Data: DesignArena Full StackOpen chartIQ BenchmarksVibe Code Bench v1.1 ScoresVals.ai Vibe Code Bench v1.1 accuracy for building web applications from scratch. Color = provider.Data: Vals.ai Vibe Code Bench, Vals.ai Vibe Code Bench v1.1Open chartIQ BenchmarksSWE Marathon ScoresEach model's SWE Marathon aggregate Resolution Rate (Pass@1). Color = provider.Data: SWE MarathonOpen chartIQ BenchmarksAPEX-SWE ScoresMercor APEX-SWE overall Pass@1 scores for economically valuable software-engineering work. Color = provider.Data: Mercor APEX-SWEOpen chartIQ BenchmarksGSO-Bench ScoresEach model's GSO Opt@1 score for software optimization tasks. Color = provider.Data: GSO-BenchOpen chartIQ BenchmarksFrontierCode Diamond ScoresEach model's FrontierCode Diamond score for high-quality production-code tasks. Color = provider.Data: Cognition FrontierCodeOpen chartIQ BenchmarksSWE-rebench Benchmark ScoresEach model's SWE-rebench resolved rate. Color = provider.Data: SWE-rebenchOpen chartIQ BenchmarksDeepSWE v1.1 Benchmark ScoresEach model's DeepSWE v1.1 Pass@1 score. Color = provider.Data: DeepSWE v1.1Open chartIQ BenchmarksKernelBench Hard Benchmark ScoresEach model's audited KernelBench Hard pass rate. Color = provider.Data: KernelBench HardOpen chartIQ BenchmarksKernelBench Mega Benchmark ScoresEach model's RTX PRO 6000 KernelBench Mega speedup over optimized PyTorch. Color = provider.Data: KernelBench MegaOpen chartIQ BenchmarksSWE-Bench Pro Benchmark ScoresEach model's SWE-Bench Pro percentage score. Color = provider.Data: Scale Labs, Scale Labs SWE-Bench Pro Public DatasetOpen chartIQ BenchmarksLiveCodeBench Benchmark ScoresEach model's LiveCodeBench score. Color = provider.Data: Vals.ai, Vals.ai LiveCodeBenchOpen chartIQ BenchmarksProgramBench ScoresAlmost Resolved score: share of ProgramBench tasks passing at least 95% of hidden tests. Color = provider.Data: Vals.ai ProgramBenchOpen chartIQ BenchmarksSciCode Benchmark ScoresEach model's SciCode score. Color = provider.Data: Artificial Analysis, Artificial Analysis Intelligence Index v4.3.2Open chartIQ BenchmarksCritPt Benchmark ScoresEach model's CritPt percentage score. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksHumanity's Last Exam Benchmark ScoresEach model's Humanity's Last Exam score. Color = provider.Data: Artificial Analysis, Artificial Analysis Intelligence Index v4.3.2Open chartIQ BenchmarksGPQA Diamond Benchmark ScoresEach model's GPQA Diamond score. Color = provider.Data: Artificial Analysis, Artificial Analysis Intelligence Index v4.3.2Open chartIQ BenchmarksMMLU-Pro Benchmark ScoresEach model's MMLU-Pro score. Color = provider.Data: Manual source capture, TIGER-Lab MMLU-Pro leaderboard (missing rows only)Open chartIQ BenchmarksMMMU-Pro Benchmark ScoresEach model's MMMU-Pro score. Color = provider.Data: Artificial Analysis, Artificial Analysis Intelligence Index v4.3.2Open chartIQ BenchmarksBrowseComp Benchmark ScoresEach model's BrowseComp score. Color = provider.Data: LLM Stats, LLM Stats BrowseCompOpen chartIQ BenchmarksOSWorld-Verified Benchmark ScoresEach model's OSWorld-Verified score. Color = provider.Data: LLM Stats, OSWorld, LLM Stats OSWorld-VerifiedOpen chartIQ BenchmarksMCP Atlas Benchmark ScoresReal-world MCP tool-use pass rates. Color = provider.Data: Scale Labs, Scale Labs MCP AtlasOpen chartIQ BenchmarksAgents' Last Exam Benchmark ScoresFull / Overall pass rates from Agents' Last Exam. Color = provider.Data: Agents' Last ExamOpen chartIQ BenchmarksArena.ai Agent Arena ScoresAggregate net improvement across real-world Agent Mode sessions. Color = provider.Data: Arena.ai Agent ArenaOpen chartIQ BenchmarksIFBench Benchmark ScoresInstruction-following and constraint-adherence scores. Color = provider.Data: Artificial Analysis, Artificial Analysis Intelligence Index v4.3.2Open chartIQ BenchmarksAA Omniscience ScoresArtificial Analysis Omniscience scores for factual reliability. Color = provider.Data: Artificial AnalysisOpen chartIQ BenchmarksBullshitBench v2 ScoresClear Pushback rate: share of attempts where a model clearly challenges a false premise instead of accepting nonsense. Color = provider.Data: BullshitBench v2Open chartIQ BenchmarksSimpleQA Verified ScoresCorrect-answer rate on Epoch AI's verified factuality set. Color = provider.Data: Epoch AI SimpleQA VerifiedOpen chartIQ BenchmarksFACTS Grounding ScoresGoogle FACTS Grounding scores for long-form responses fully grounded in provided documents. Color = provider.Data: Kaggle FACTS Grounding leaderboardOpen chartIQ BenchmarksIOI V2 Scores2024–2026 olympiad tasks, no interactive grading feedback. Version 2 only; separately calibrated from archived V1. Contributes to Programmatic Reasoning using a provisional calibrated scale.Data: Vals.ai IOI V2Open chartIQ BenchmarksProofBench 1.1 ScoresNative Lean 4.25.2 verification including native_decide. Version 1.1 only; not comparable to the unversioned historical cohort. Contributes to Mathematical Reasoning using a provisional calibrated scale.Data: ProofBench 1.1Open chartIQ BenchmarksTerminal-Bench 4.0 ScoresExact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Programmatic Reasoning using a provisional calibrated scale.Data: Artificial AnalysisOpen chartIQ BenchmarksAA Long Context Reasoning v1.1 ScoresExact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Reliability using a provisional calibrated scale.Data: Artificial AnalysisOpen chartIQ BenchmarksGDPval-AA v2.1 normalized score ScoresExact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Knowledge Work using a provisional calibrated scale.Data: Artificial AnalysisOpen chartIQ BenchmarksAutomationBench-AA ScoresExact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Knowledge Work using a provisional calibrated scale.Data: Artificial AnalysisOpen chartIQ BenchmarksGDP.pdf ScoresExact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Knowledge Work using a provisional calibrated scale.Data: Artificial AnalysisOpen chartIQ BenchmarksAA-Briefcase v1.1 Elo ScoresExact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Knowledge Work using a provisional calibrated scale.Data: Artificial AnalysisOpen chartIQ BenchmarksMirrorCode ScoresEpoch ML,+Private,2L: 30 tasks, three repeats, seven days and 10B tokens per attempt. Average rate of completely solved programs; distinct from the paper cohort. Contributes to Programmatic Reasoning using a provisional calibrated scale.Data: Epoch AIOpen chartIQ BenchmarksAA-AnalystAgent ScoresQuantitative analysis from spreadsheets and documents. Headline pass^5: correct on every one of five attempts, not pass@5. Contributes to Knowledge Work using a provisional calibrated scale.Data: Artificial AnalysisOpen chartIQ BenchmarksEnterpriseOps-Gym-AA ScoresStrict success on 1,117 oracle-mode enterprise workflows, graded against final database state using AA’s Stirrup harness. Contributes to Knowledge Work using a provisional calibrated scale.Data: Artificial AnalysisOpen chartIQ BenchmarksChess Puzzles ScoresEpoch’s 100 novel chess positions, presented as text (FEN). Exact best-next-move accuracy; no vision input. Chess knowledge and planning both contribute. Contributes to Abstract Reasoning using a provisional calibrated scale.Data: Epoch AIOpen chartIQ BenchmarksMystery Game Puzzles ScoresEpoch’s 100 text-state puzzles in an undisclosed game. Best-next-move accuracy in a minimal agent scaffold; invalid or missing moves receive no credit. Prompts and transcripts are private. Contributes to Abstract Reasoning using a provisional calibrated scale.Data: Epoch AIOpen chart
Emotional & Interaction
Emotional & InteractionEmotional Reasoning IQ Bell CurveEach model's Emotional Reasoning IQ, derived from EQ-Bench 4 and AttuneBench and excluded from Composite IQ, placed on a standard normal IQ distribution.Data: EQ-Bench 4 Elo, EQ-Bench 3, AttuneBenchOpen chartEmotional & InteractionArena.ai Overall Benchmark ScoresHuman-preference Arena.ai Overall Elo scores. Color = provider.Data: Arena.aiOpen chartEmotional & InteractionEQ-Bench 3 Benchmark ScoresAI-judged EQ-Bench 3 Elo, a style-sensitive emotional reasoning signal. Color = provider.Data: EQ-Bench 3Open chartEmotional & InteractionAttuneBench Benchmark ScoresDefault-mode AttuneBench Composite scores from real human-AI conversations. Color = provider.Data: AttuneBenchOpen chartEmotional & InteractionEQ-Bench 4 Elo ScoresPublished multi-judge Elo on 120 personas and 16 turns. No inherited EQ-Bench 3 Anthropic penalty. Subjective roleplay evaluation; not a validated human EI measure. Contributes to Emotional Reasoning using a provisional calibrated scale.Data: EQ-Bench 4 EloOpen chart
Cost
CostIQ vs Input/Output Token CostEach model's estimated IQ plotted against its published token price. Toggle between input and output price per 1M tokens.Data: AI IQ methodologyOpen chartCostIQ vs Blended Token CostEach model's estimated IQ plotted against the cost of a representative 1M-token blend. Toggle between a coding blend (cache-heavy — 800K cache-read + 100K input + 100K output) and a copywriting blend (output-heavy — 150K input + 850K output).Data: AI IQ methodologyOpen chartCostIQ vs CostEach model's estimated IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: AI IQ methodologyOpen chartCostThe Falling Cost of IntelligenceX = release date. Y = effective cost per 1M I/O Tokens (log). Each step line tracks the cheapest model to date at or above an IQ level; large dots mark the models that set each new price floor, faint dots show every other qualifying model, colored by provider.Data: AI IQ methodologyOpen chartCostEQ vs Effective CostDiagnostic EQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +4 moreOpen chartCostAI Models by CostPublished price for 1M I/O Tokens compared with effective cost after applying the measured or imputed task-usage multiplier.Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartCostInput Price vs Output PriceEach model positioned by its published token prices — input price (Y) against output price (X), both per 1M tokens on a log scale. The dashed line is the best fit across models, showing the typical output-to-input price relationship.Data: AI IQ datasetOpen chartCostTask EfficiencyEach dot shows the inverse of the effective-cost usage multiplier. Higher means less price-adjusted task work: 2× is about half the median task effort. Source-backed multipliers are preferred; lineage, peer, and 1× fallbacks are labeled in tooltips.Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartCostSticker Price vs Effective CostEach model's published sticker price for 1M I/O Tokens plotted against its effective cost after the measured or imputed task-usage multiplier. Points above the dashed line cost more per task than their sticker price suggests.Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chart
Response Time
Response TimeIQ vs ThroughputThroughput is median output tokens per second (TPS). Each model's estimated IQ is plotted against how quickly it streams its answer after generation begins.Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeIQ vs LatencyLatency is time to first token (TTFT): how long a model takes to begin responding. Drag the input-length slider from 1K to 10K tokens to interpolate between AA's two measured TTFT anchors.Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeIQ vs SpeedEach model's estimated IQ against its completion time, which is time to last token (TTLT). Drag the input (1K–10K tokens) and output (100–10K tokens) sliders to see how it changes with the workload.Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeIQ vs Completion Time vs Cost in 3D3D scatter: X = completion time (time to last token, TTLT; log, faster to the right), Y = IQ, Z = effective cost (log). Color = provider. Drag to rotate.Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartResponse TimeAnimating Completion Time over IQ vs CostX = Effective Cost/1M I/O Tokens (log). Y = IQ. Clock spin rate represents completion time (time to last token, TTLT).Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeAnimating Cost over IQ vs Completion TimeX = completion time (time to last token, TTLT; log, reversed). Y = IQ. Effective Cost/1M I/O Tokens is shown as bills over a 10s period.Data: AI IQ methodology, Artificial AnalysisOpen chart
Applied Domains
Applied DomainsEngineering Design IQ Bell CurveEach model's Engineering Design IQ, placed on a standard normal IQ distribution. Domain composites are provisional and sit outside Composite IQ.Data: AI IQ datasetOpen chartApplied DomainsEngineering Design IQ vs CostEach model's Engineering Design IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartApplied DomainsEngineering Design IQ RankingsEach model's estimated Engineering Design IQ, ranked highest to lowest. Defaults to current-generation models.Data: AI IQ datasetOpen chartApplied DomainsEngineering Design IQ vs SpeedEach model's estimated Engineering Design IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial AnalysisOpen chartApplied DomainsEngineering Design IQ Over TimeEach model's estimated Engineering Design IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartApplied DomainsCybersecurity IQ Bell CurveEach model's Cybersecurity IQ, placed on a standard normal IQ distribution. Domain composites are provisional and sit outside Composite IQ.Data: AI IQ datasetOpen chartApplied DomainsCybersecurity IQ vs CostEach model's Cybersecurity IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartApplied DomainsCybersecurity IQ RankingsEach model's estimated Cybersecurity IQ, ranked highest to lowest. Defaults to current-generation models.Data: AI IQ datasetOpen chartApplied DomainsCybersecurity IQ vs SpeedEach model's estimated Cybersecurity IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial AnalysisOpen chartApplied DomainsCybersecurity IQ Over TimeEach model's estimated Cybersecurity IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartApplied DomainsBiology IQ Bell CurveEach model's Biology IQ, placed on a standard normal IQ distribution. Domain composites are provisional and sit outside Composite IQ.Data: AI IQ datasetOpen chartApplied DomainsBiology IQ vs CostEach model's Biology IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartApplied DomainsBiology IQ RankingsEach model's estimated Biology IQ, ranked highest to lowest. Defaults to current-generation models.Data: AI IQ datasetOpen chartApplied DomainsBiology IQ vs SpeedEach model's estimated Biology IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial AnalysisOpen chartApplied DomainsBiology IQ Over TimeEach model's estimated Biology IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartApplied DomainsAI/ML Engineering IQ Bell CurveEach model's AI/ML Engineering IQ, placed on a standard normal IQ distribution. Domain composites are provisional and sit outside Composite IQ.Data: AI IQ datasetOpen chartApplied DomainsAI/ML Engineering IQ vs CostEach model's AI/ML Engineering IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartApplied DomainsAI/ML Engineering IQ RankingsEach model's estimated AI/ML Engineering IQ, ranked highest to lowest. Defaults to current-generation models.Data: AI IQ datasetOpen chartApplied DomainsAI/ML Engineering IQ vs SpeedEach model's estimated AI/ML Engineering IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial AnalysisOpen chartApplied DomainsAI/ML Engineering IQ Over TimeEach model's estimated AI/ML Engineering IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartApplied DomainsFinance IQ Bell CurveEach model's Finance IQ, placed on a standard normal IQ distribution. Domain composites are provisional and sit outside Composite IQ.Data: AI IQ datasetOpen chartApplied DomainsFinance IQ vs CostEach model's Finance IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartApplied DomainsFinance IQ RankingsEach model's estimated Finance IQ, ranked highest to lowest. Defaults to current-generation models.Data: AI IQ datasetOpen chartApplied DomainsFinance IQ vs SpeedEach model's estimated Finance IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial AnalysisOpen chartApplied DomainsFinance IQ Over TimeEach model's estimated Finance IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartApplied DomainsWriting IQ Bell CurveEach model's Writing IQ, placed on a standard normal IQ distribution. Domain composites are provisional and sit outside Composite IQ.Data: AI IQ datasetOpen chartApplied DomainsWriting IQ vs CostEach model's Writing IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartApplied DomainsWriting IQ RankingsEach model's estimated Writing IQ, ranked highest to lowest. Defaults to current-generation models.Data: AI IQ datasetOpen chartApplied DomainsWriting IQ vs SpeedEach model's estimated Writing IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial AnalysisOpen chartApplied DomainsWriting IQ Over TimeEach model's estimated Writing IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartApplied DomainsSoftware Engineering IQ Bell CurveEach model's Software Engineering IQ, placed on a standard normal IQ distribution. Domain composites are provisional and sit outside Composite IQ.Data: AI IQ datasetOpen chartApplied DomainsSoftware Engineering IQ vs CostEach model's Software Engineering IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartApplied DomainsSoftware Engineering IQ RankingsEach model's estimated Software Engineering IQ, ranked highest to lowest. Defaults to current-generation models.Data: AI IQ datasetOpen chartApplied DomainsSoftware Engineering IQ vs SpeedEach model's estimated Software Engineering IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial AnalysisOpen chartApplied DomainsSoftware Engineering IQ Over TimeEach model's estimated Software Engineering IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartApplied DomainsWebDev & Design IQ Bell CurveEach model's WebDev & Design IQ, placed on a standard normal IQ distribution. Domain composites are provisional and sit outside Composite IQ.Data: AI IQ datasetOpen chartApplied DomainsWebDev & Design IQ vs CostEach model's WebDev & Design IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartApplied DomainsWebDev & Design IQ RankingsEach model's estimated WebDev & Design IQ, ranked highest to lowest. Defaults to current-generation models.Data: AI IQ datasetOpen chartApplied DomainsWebDev & Design IQ vs SpeedEach model's estimated WebDev & Design IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial AnalysisOpen chartApplied DomainsWebDev & Design IQ Over TimeEach model's estimated WebDev & Design IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartApplied DomainsEmotional Reasoning IQ RankingsEach model's estimated Emotional Reasoning IQ, ranked highest to lowest. Defaults to current-generation models.Data: EQ-Bench 4 Elo, EQ-Bench 3, AttuneBenchOpen chartApplied DomainsEmotional Reasoning IQ vs SpeedEach model's estimated Emotional Reasoning IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: EQ-Bench 4 Elo, EQ-Bench 3, AttuneBench +1 moreOpen chartApplied DomainsEmotional Reasoning IQ Over TimeEach model's estimated Emotional Reasoning IQ by release date. Step-lines connect each provider's flagship models over time.Data: EQ-Bench 4 Elo, EQ-Bench 3, AttuneBench +29 moreOpen chartApplied DomainsLegal IQ Bell CurveEach model's Legal IQ, placed on a standard normal IQ distribution. Domain composites are provisional and sit outside Composite IQ.Data: AI IQ datasetOpen chartApplied DomainsLegal IQ vs CostEach model's Legal IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartApplied DomainsLegal IQ RankingsEach model's estimated Legal IQ, ranked highest to lowest. Defaults to current-generation models.Data: AI IQ datasetOpen chartApplied DomainsLegal IQ vs SpeedEach model's estimated Legal IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial AnalysisOpen chartApplied DomainsLegal IQ Over TimeEach model's estimated Legal IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chartApplied DomainsHealthcare IQ Bell CurveEach model's Healthcare IQ, placed on a standard normal IQ distribution. Domain composites are provisional and sit outside Composite IQ.Data: AI IQ datasetOpen chartApplied DomainsHealthcare IQ vs CostEach model's Healthcare IQ against its effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: Artificial Analysis, ARC Prize, Vals Index v2 Cost per Test +1 moreOpen chartApplied DomainsHealthcare IQ RankingsEach model's estimated Healthcare IQ, ranked highest to lowest. Defaults to current-generation models.Data: AI IQ datasetOpen chartApplied DomainsHealthcare IQ vs SpeedEach model's estimated Healthcare IQ against its completion time, which is time to last token (TTLT). Drag the input and output sliders to see how it changes with the workload.Data: Artificial AnalysisOpen chartApplied DomainsHealthcare IQ Over TimeEach model's estimated Healthcare IQ by release date. Step-lines connect each provider's flagship models over time.Data: ARC Prize, Epoch AI, Epoch AI +26 moreOpen chart
Diagnostic Benchmarks
Diagnostic BenchmarksMultiChallenge ScoresDiagnostic multi-turn instruction retention, inference memory, editing, and self-coherence. Does not contribute to IQ. Color = provider.Data: Scale Labs MultiChallengeOpen chartDiagnostic BenchmarksTerminal-Bench Science 0.1 (Vals) ScoresAll 70 tasks from v0.1.0; fixed Terminus 2 harness, pass@1, eight-hour budget. Infrastructure retries permitted. Not interchangeable with native-agent official scores. Standalone raw result; does not contribute to IQ/EQ.Data: Terminal-Bench Science 0.1 (Vals)Open chartDiagnostic BenchmarksTax Agent Bench ScoresWeighted partial credit over the private 193-question test set, with must-pass checks and citation adjustment. Not the secondary All-Pass metric. Standalone raw result; does not contribute to IQ/EQ.Data: Tax Agent BenchOpen chartDiagnostic BenchmarksVals Index v2 ScoresVersion 2 cohort only. Accuracy, cost and latency must come from the same source configuration. Highest published effort; Max and default retained separately in source snapshot. Standalone raw result; does not contribute to IQ/EQ.Data: Vals Index v2Open chartDiagnostic BenchmarksArXivMath August 2026 ScoresAugust 2026 release under the September 15 grading revision. Agentic offline harness. Models released after the problems carry source contamination warnings; interpret as diagnostic evidence. Standalone raw result; does not contribute to IQ/EQ.Data: ArXivMath August 2026Open chartDiagnostic BenchmarksBrokenArXiv August 2026 ScoresAugust 2026 release under the September 15 grading revision. Agentic offline harness. Models released after the problems carry source contamination warnings; interpret as diagnostic evidence. Standalone raw result; does not contribute to IQ/EQ.Data: BrokenArXiv August 2026Open chartDiagnostic BenchmarksOSWorld 2.0 ScoresOfficial OSWorld 2.0 leaderboard only: 108 long-horizon desktop tasks, 500-step budget, full task set, binary completion accuracy, best reasoning/tool configuration per model. Partial-credit scores and provider-reported launch figures are excluded. Standalone raw result; does not contribute to IQ/EQ.Data: OSWorld 2.0 official leaderboardOpen chart