Speed

How models trade off intelligence against speed

IQ vs Latency (Time to First Token)
Each model's estimated IQ plotted against its time to first token (TTFT). Drag the input-length slider from 1K to 10K tokens — dots glide between AA's two measured anchors (values in between are linearly interpolated).
Controls:
Input10K tok
IQ1:1Speed

How to read this chart

Each model's estimated IQ plotted against its time to first token (TTFT). Drag the input-length slider from 1K to 10K tokens — dots glide between AA's two measured anchors (values in between are linearly interpolated).

IQ vs Throughput (Tokens Per Second)
Each model's estimated IQ plotted against its median output speed in tokens per second
Controls:
IQ1:1Speed

How to read this chart

Each point is a public model. The chart compares IQ against Tokens/s, with color showing the model provider.

IQ vs End-to-End Response Time
Each model's IQ against its end-to-end response time, built from its own components: time to first answer token (which grows with input length) plus output tokens ÷ tokens-per-second. Drag the input slider (1K–10K tokens) and output slider (100–10K tokens) to see how total latency scales.
Controls:
Input10K tok
Output1K tok
IQ1:1Speed

How to read this chart

Each model's IQ against its end-to-end response time, built from its own components: time to first answer token (which grows with input length) plus output tokens ÷ tokens-per-second. Drag the input slider (1K–10K tokens) and output slider (100–10K tokens) to see how total latency scales.