Speed
How models trade off intelligence against speed
IQ vs Latency (Time to First Token)
IQ vs Latency (Time to First Token)
Each model's estimated IQ plotted against its time to first token (TTFT). Drag the input-length slider from 1K to 10K tokens — dots glide between AA's two measured anchors (values in between are linearly interpolated).
Controls:
How to read this chart
Each model's estimated IQ plotted against its time to first token (TTFT). Drag the input-length slider from 1K to 10K tokens — dots glide between AA's two measured anchors (values in between are linearly interpolated).
Data sources
IQ vs Throughput (Tokens Per Second)
IQ vs Throughput (Tokens Per Second)
Each model's estimated IQ plotted against its median output speed in tokens per second
Controls:
How to read this chart
Each point is a public model. The chart compares IQ against Tokens/s, with color showing the model provider.
Data sources
IQ vs End-to-End Response Time
IQ vs End-to-End Response Time
Each model's IQ against its end-to-end response time, built from its own components: time to first answer token (which grows with input length) plus output tokens ÷ tokens-per-second. Drag the input slider (1K–10K tokens) and output slider (100–10K tokens) to see how total latency scales.
Controls:
How to read this chart
Each model's IQ against its end-to-end response time, built from its own components: time to first answer token (which grows with input length) plus output tokens ÷ tokens-per-second. Drag the input slider (1K–10K tokens) and output slider (100–10K tokens) to see how total latency scales.
Data sources