IQ vs Latency
Latency is time to first token (TTFT): how long a model takes to begin responding. Drag the input-length slider from 1K to 10K tokens to interpolate between AA's two measured TTFT anchors.
IQ vs Latency
IQ vs Latency
Latency is time to first token (TTFT): how long a model takes to begin responding. Drag the input-length slider from 1K to 10K tokens to interpolate between AA's two measured TTFT anchors.
Controls:
How to read this chart
Latency is time to first token (TTFT): how long a model takes to begin responding. Drag the input-length slider from 1K to 10K tokens to interpolate between AA's two measured TTFT anchors.
Data sources
Response TimeIQ vs Completion TimeCompletion time is time to last token (TTLT): time to first answer token plus output tokens ÷ throughput (TPS). Drag the input slider (1K–10K tokens) and output slider (100–10K tokens) to see how TTLT changes with the workload.Data: AI IQ methodology, Artificial AnalysisOpen chartResponse TimeIQ vs ThroughputThroughput is median output tokens per second (TPS). Each model's estimated IQ is plotted against how quickly it streams its answer after generation begins.Data: AI IQ methodology, Artificial AnalysisOpen chartCostIQ vs Effective CostEach model's estimated IQ plotted against effective cost per 1M I/O Tokens (sticker price × measured or imputed usage multiplier).Data: AI IQ methodologyOpen chartCoreAI Models on the IQ Bell CurveEach model's estimated IQ plotted on a standard normal IQ distributionData: ARC Prize, Epoch AI FrontierMath, Vals.ai +24 moreOpen chartCoreIQ vs EQ vs Cost in 3D3D scatter: X = diagnostic EQ, Y = IQ, Z = effective cost (log). Color = provider. Drag to rotate.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chartResponse TimeIQ vs Completion Time vs Cost in 3D3D scatter: X = completion time (time to last token, TTLT; log, faster to the right), Y = IQ, Z = effective cost (log). Color = provider. Drag to rotate.Data: Artificial Analysis, ARC Prize, Vals.aiOpen chart