IQ vs End-to-End Response Time

Each model's IQ against its end-to-end response time, built from its own components: time to first answer token (which grows with input length) plus output tokens ÷ tokens-per-second. Drag the input slider (1K–10K tokens) and output slider (100–10K tokens) to see how total latency scales.

IQ vs End-to-End Response Time
Each model's IQ against its end-to-end response time, built from its own components: time to first answer token (which grows with input length) plus output tokens ÷ tokens-per-second. Drag the input slider (1K–10K tokens) and output slider (100–10K tokens) to see how total latency scales.
Controls:
Input10K tok
Output1K tok
IQ1:1Speed

How to read this chart

Each model's IQ against its end-to-end response time, built from its own components: time to first answer token (which grows with input length) plus output tokens ÷ tokens-per-second. Drag the input slider (1K–10K tokens) and output slider (100–10K tokens) to see how total latency scales.