About
Independent, transparent measurement of AI model intelligence
Our mission
AI IQ exists to answer a deceptively hard question: how capable are AI models, really? Benchmark results are scattered across papers, leaderboards, and vendor announcements — each with its own scale, its own blind spots, and its own incentives. We consolidate them into a single, rigorous, human-readable measure of model intelligence, so that engineers, researchers, and decision-makers can compare models on evidence rather than marketing.
What we publish
AI IQ scores frontier models on a human IQ scale, built from seven scored dimensions spanning abstract reasoning, mathematics, science, frontend and backend engineering, computer use, and reliability. Alongside the headline rankings, we publish per-dimension breakdowns, cost and speed tradeoffs measured in real-world terms, domain-specific leaderboards for fields like biotechnology and cybersecurity, and a free public JSON API so the data can be used anywhere.
How we work
Three principles govern every number on this site:
Source-backed. Every score derives from published benchmark results, and our data pipeline tracks where each result came from. We do not invent measurements, and where coverage is missing we say so rather than paper over it.
Transparent. The full scoring approach — how raw benchmark scores map to IQ values, how dimensions are averaged, and how missing data is conservatively handled — is documented in our public Methodology. If you can't check our math, you shouldn't trust our rankings.
Independent. AI IQ is not affiliated with, funded by, or influenced by any AI lab or model vendor. Placement on our leaderboards cannot be bought, and models from every provider are held to the same standard.
Stay in the loop
Model rankings shift quickly. The AI IQ newsletter covers new model releases, ranking changes, and what they mean — subscribe below, or build with the API.
Origins
AI IQ was founded by Ryan Shea, an engineer and repeat founder. It began as a private tool for tracking which models were actually getting better — and grew into the public rankings, tools, and API it is today.