AA Long Context Reasoning v1.1

A bird's-eye view of the benchmark: what it measures and every AI IQ chart built on it.

About AA Long Context Reasoning v1.1

AA Long Context Reasoning v1.1 has source-backed multi-effort configurations. The Cost Efficiency chart expands only those configurations; the score bar remains one canonical row per model.

AA Long Context Reasoning v1.1 Cost Efficiency
Source-matched reasoning-effort score and task-cost configurations. Each line is one canonical model. Color = provider.

How to read this chart

Each line uses the shared model-configuration dataset to compare source-matched reasoning-effort levels on AA Long Context Reasoning v1.1. Every point shows the score and AA cost per task for that exact configuration. Canonical score bars and bell curves remain one point per model.

AA Long Context Reasoning v1.1 Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Reliability using a provisional calibrated scale.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Sep 22, 2026
AA Long Context Reasoning v1.1 vs Effective Cost
X = effective cost (log). Y = AA Long Context Reasoning v1.1. Exact AA v4.3.2 model/configuration; never substitute older benchmark versions.
Controls:
AA Long Context Reasoning v1.11:1Cost

How to read this chart

Each point is a public model. The chart compares AA Long Context Reasoning v1.1 (%) against Effective Cost (per 1M I/O Tokens), with color showing the model provider.