EnterpriseOps-Gym-AA

A bird's-eye view of the benchmark: what it measures and every AI IQ chart built on it.

EnterpriseOps-Gym-AA Cost Efficiency
X-axis
EnterpriseOps-Gym-AA Scores
EnterpriseOps-Gym-AA Model Table
#ModelEffortScoreEffective cost (est.)Completion time (est.)Weights

About EnterpriseOps-Gym-AA

What EnterpriseOps-Gym-AA measures

EnterpriseOps-Gym-AA does not yet have two source-matched reasoning-effort score/cost rows for one model, so the Cost Efficiency chart shows one point per model and will connect effort levels when qualifying data is imported.

Reading the cost view

EnterpriseOps-Gym-AA publishes no per-run cost, so this chart falls back to each model's estimated effective cost per 1M I/O tokens (a runtime estimate from measured or imputed token usage, not a published task cost). Explicit reasoning/nonreasoning siblings are connected where both are public; every other model is a standalone point. Up and to the left is better.

Reading the completion-time view

Each point is a public model's EnterpriseOps-Gym-AA score against its estimated completion time (time to last token) for the input and output lengths set by the sliders. Timing comes from the site's per-model speed measurements, not from the benchmark run, so every point is an estimate and labels carry a ~ prefix. The newest generation with a published result is shown by default. Up and to the left is better.

Reading the score bars

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Reading the model table

Each row is one canonical model: the score shown is the same best published configuration the score bars use. Models with several published reasoning-effort levels expand to one sub-row per level, with the score and effective cost per 1M I/O Tokens the source reports for that level. This benchmark publishes no per-run cost, so the cost column is the site's estimated effective cost per 1M I/O tokens (marked ~). Completion time is the site's estimated time to last token for a 10K-token input and 1K-token output (marked ~), not a benchmark run time. The top 20 rows are shown by default; click a column heading to sort, and missing values always sort last.