# AI IQ > Intelligently measuring AI intelligence. Compare AI models across Composite IQ, its six scored dimensions, applied-capability domains, speed, and cost. AI IQ (aiiq.org) is an open data project by Liberated Software LLC that assigns Composite IQ scores to AI models using public benchmarks across six scored capability dimensions, then visualizes applied-capability domains and tradeoffs against cost and speed. Manual source captures are archived separately from extracted values so benchmark updates can be re-parsed later. ## Data API Public API documentation is available at: - [API Documentation](https://www.aiiq.org/docs/api/): Endpoint documentation, example responses, and privacy boundary. - **Canonical base URL**: https://www.aiiq.org/api/v1/ (permanent alias; bare https://www.aiiq.org/api/ also works) - **OpenAPI 3.1 spec**: https://www.aiiq.org/api/openapi.json (machine-readable endpoint and schema definitions) Every endpoint returns a versioned envelope: `{ "apiVersion": "1", "methodologyVersion": "...", "updatedAt": "...", }`. Collection endpoints place their array under a named key (`models`, `benchmarks`, `rankings`, `charts`); the model-detail endpoint places model fields at the top level alongside the envelope keys. Structured JSON data is available at: - [Models](https://www.aiiq.org/api/models): Sanitized public model summaries with stable IDs, rank, composite IQ, a per-dimension IQ map, a cost block (input/output/effective cost per 1M tokens), timestamp, and canonical URL. - [Model Detail](https://www.aiiq.org/api/models/gpt-5.5): Detailed public model data with full per-dimension blocks (IQ + coverage), the cost block, source-backed benchmark results, and methodology version. - [Benchmarks](https://www.aiiq.org/api/benchmarks): Benchmark metadata including the dimension each benchmark feeds, direction, and unit. - [Rankings](https://www.aiiq.org/api/rankings): Ranking catalog (Composite IQ, Effective Cost, six per-dimension IQ rankings, per-benchmark rankings) with model counts and detail URLs; fetch one full leaderboard from /api/rankings/:id. - [Charts](https://www.aiiq.org/api/charts): Canonical chart metadata. - [Methodology](https://www.aiiq.org/api/methodology): Public methodology metadata. - [Domains](https://www.aiiq.org/api/domains): Applied-capability domains tracked outside Composite IQ: Cybersecurity, Biology, Machine Learning Engineering, Writing, Software Engineering, WebDev & Design, and Emotional Reasoning (EQ). `/api/domains/:slug` returns per-model domain composite IQs plus the domain's benchmark leaderboards with model + harness rows. ### Schema Each model summary object contains: | Field | Type | Description | |---|---|---| | id | string | Stable model identifier, e.g. "gpt-5.5" | | name | string | Human-readable model name | | provider | string | Public provider grouping | | rank | number\|null | Current Composite IQ rank when ranked | | iq | number\|null | Rounded Composite IQ | | dimensions | object | Map of the six scored dimension slugs to rounded dimension IQ (`null` when a dimension is unscored) | | emotionalReasoning | number\|null | Emotional Reasoning (EQ) domain score, excluded from Composite IQ | | cost | object | Cost block: `inputPer1M`, `outputPer1M`, `ioPer1M`, `usageMultiplier`, `effectivePer1M`, `unit` | | updatedAt | string | ISO update timestamp | | url | string | Canonical public model profile URL | The six scored dimension slugs are `abstract-reasoning`, `mathematical-reasoning`, `academic-reasoning`, `programmatic-reasoning`, `computer-use`, and `reliability`. Emotional Reasoning (EQ) is a separate domain; the standalone `composite-eq` ranking has been retired. ## IQ Methodology IQ scores are computed from benchmarks organized into six equally weighted capability dimensions: - **Abstract Reasoning**: ARC-AGI-3, ARC-AGI-2, ARC-AGI-1 - **Mathematical Reasoning**: FrontierMath Tier 4, FrontierMath Tier 1-3, ProofBench 1.1, MathArena, AIME (transitional) - **Academic Reasoning**: Humanity's Last Exam, GPQA Diamond, CritPt, SciCode, MMLU-Pro, MMMU-Pro - **Programmatic Reasoning**: LiveCodeBench (historical coverage), IOI V2, Terminal-Bench 4.0, ProgramBench (scored by Almost Resolved), SWE-rebench - **Computer Use**: BrowseComp, OSWorld-Verified, MCP Atlas, Arena.ai Agent Arena, Agents' Last Exam - **Reliability**: SimpleQA Verified, AA Omniscience, BullshitBench v2, IFBench, AA Long Context Reasoning v1.1, FACTS Grounding Each benchmark score is mapped to an IQ-equivalent via piecewise-linear interpolation through an expected-score ladder. Every scored benchmark defines the source score expected at IQ 70, 85, 100, 115, 130, 145, and 160; saturated or gameable benchmarks are constrained by the shape of that ladder rather than by a separate cap. ARC-AGI-1/2/3 are jointly calibrated: once a model has any direct ARC result, Abstract IQ is the equal mean of only its available source-backed ARC projections, so later coverage raises confidence and produces an order-invariant update. Models with zero direct ARC coverage can still use the conservative scoring-only ARC lineage and peer fallbacks. D2–D6 use Bayesian latent dimension ability with benchmark offsets and one uncertain predecessor prior. Only direct observations enter the likelihood. Rankings use posterior means; model-detail dimensions expose conditional, uncalibrated uncertainty ranges separately. Whole-dimension cold-start benchmark prediction is a known weakness relative to direct predecessor carry-forward. Abstract Reasoning retains its ARC-specific policy. FrontierSWE V1 and IOI V1 are archived; IOI V2 is independently calibrated. OSWorld 2.0 and MultiChallenge are raw diagnostics outside Composite IQ. Derived IQ uses six equally weighted dimensions. Whole dimensions with neither direct nor compatible ancestral evidence use a matched lower-quartile waterfall: comparable models with real data for the missing dimension and similar other-dimension capability, then nearest comparable models, then the global dimension lower quartile. If no real data exists for that dimension, no neutral default is invented and the model does not receive a derived all-dimension IQ. Emotional Reasoning (EQ) is a separate domain computed from EQ-Bench 4, the frozen EQ-Bench 3 cohort, and AttuneBench and excluded from Composite IQ. Raw benchmark fields in the public dataset remain source-backed. Imputed benchmark and dimension values are used for derived IQ scoring only; composite IQ, dimension IQs, and effective cost are runtime-derived values. ## Pages - [Home](https://www.aiiq.org/): Interactive charts — IQ bell curve, IQ vs cost, IQ vs speed, benchmark comparisons - [Methodology](https://www.aiiq.org/methodology/): Full scoring methodology with IQ 70-160 expected-score ladders and math - [API](https://www.aiiq.org/docs/api/): Public API documentation - [Models](https://www.aiiq.org/models/): Public model directory with links to model profile pages - [Model API Pricing](https://www.aiiq.org/models/gpt-5.5/pricing/): Every model with published API prices has a pricing page at `/models//pricing/` — input/output/cache prices, workload cost examples, effective cost, and cheaper alternatives at similar capability. - [Compare](https://www.aiiq.org/compare/): Side-by-side model comparison hub. Any pair of public models has a page at `/compare/-vs-/` with the two ids in alphabetical order (e.g. https://www.aiiq.org/compare/fable-5-vs-gpt-5.6-sol/), comparing rank, Composite IQ, the six dimension IQs, Emotional Reasoning, effective cost, sticker price, speed, and context window. Each page embeds the underlying numbers (dimension IQs and applied-domain composite IQs for both models) as JSON in a `