Speed
How models trade off intelligence against speed
From prompt to finished answer
First comes the wait for the first token, then the answer streams out. The blue rectangle is that stream: its area is the output length, its height is the throughput, so its width is the streaming time. Wait plus stream spans the completion time.
How to read this chart
Each model's estimated IQ against its completion time, which is time to last token (TTLT). Drag the input (1K–10K tokens) and output (100–10K tokens) sliders to see how it changes with the workload.
How to read this chart
Latency is time to first token (TTFT): how long a model takes to begin responding. Drag the input-length slider from 1K to 10K tokens to interpolate between AA's two measured TTFT anchors.
How to read this chart
Each point is a public model. The chart compares IQ against Throughput, with color showing the model provider.