AA-Briefcase v1.1 Elo Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Raw-only until separately calibrated for AI IQ. Standalone raw result; does not contribute to IQ/EQ.
AA-Briefcase v1.1 Elo Scores
AA-Briefcase v1.1 Elo Scores
Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Raw-only until separately calibrated for AI IQ. Standalone raw result; does not contribute to IQ/EQ.
How to read this chart
Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.
AA-Briefcase v1.1 Elo vs Effective Cost
AA-Briefcase v1.1 Elo vs Effective Cost
X = effective cost (log). Y = AA-Briefcase v1.1 Elo. Exact AA v4.3.2 model/configuration; never substitute older benchmark versions. Raw-only until separately calibrated for AI IQ.
Controls:
How to read this chart
Each point is a public model. The chart compares AA-Briefcase v1.1 Elo against Effective Cost (per 1M I/O Tokens), with color showing the model provider.
Diagnostic BenchmarksTerminal-Bench Science 0.1 (Vals) ScoresAll 70 tasks from v0.1.0; fixed Terminus 2 harness, pass@1, eight-hour budget. Infrastructure retries permitted. Not interchangeable with native-agent official scores. Standalone raw result; does not contribute to IQ/EQ.Data: Terminal-Bench Science 0.1 (Vals)Open chartDiagnostic BenchmarksTax Agent Bench ScoresWeighted partial credit over the private 193-question test set, with must-pass checks and citation adjustment. Not the secondary All-Pass metric. Standalone raw result; does not contribute to IQ/EQ.Data: Tax Agent BenchOpen chartDiagnostic BenchmarksVals Index v2 ScoresVersion 2 cohort only. Accuracy, cost and latency must come from the same source configuration. Highest published effort; Max and default retained separately in source snapshot. Standalone raw result; does not contribute to IQ/EQ.Data: Vals Index v2Open chartDiagnostic BenchmarksArXivMath August 2026 ScoresAugust 2026 release under the September 15 grading revision. Agentic offline harness. Models released after the problems carry source contamination warnings; interpret as diagnostic evidence. Standalone raw result; does not contribute to IQ/EQ.Data: ArXivMath August 2026Open chartDiagnostic BenchmarksBrokenArXiv August 2026 ScoresAugust 2026 release under the September 15 grading revision. Agentic offline harness. Models released after the problems carry source contamination warnings; interpret as diagnostic evidence. Standalone raw result; does not contribute to IQ/EQ.Data: BrokenArXiv August 2026Open chartDiagnostic BenchmarksGDPval-AA v2.1 normalized score ScoresExact AA v4.3.2 model/configuration; never substitute older benchmark versions. Raw-only until separately calibrated for AI IQ. Standalone raw result; does not contribute to IQ/EQ.Data: Artificial AnalysisOpen chart