AA Long Context Reasoning v1.1 Scores

Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Reliability using a provisional calibrated scale.

AA Long Context Reasoning v1.1 Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Contributes to Reliability using a provisional calibrated scale.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Sep 22, 2026
AA Long Context Reasoning v1.1 vs Effective Cost
X = effective cost (log). Y = AA Long Context Reasoning v1.1. Exact AA v4.3.2 model/configuration; never substitute older benchmark versions.
Controls:
AA Long Context Reasoning v1.11:1Cost

How to read this chart

Each point is a public model. The chart compares AA Long Context Reasoning v1.1 (%) against Effective Cost (per 1M I/O Tokens), with color showing the model provider.