Terminal-Bench Science 0.1 (Vals) Scores

All 70 tasks from v0.1.0; fixed Terminus 2 harness, pass@1, eight-hour budget. Infrastructure retries permitted. Not interchangeable with native-agent official scores. Standalone raw result; does not contribute to IQ/EQ.

Terminal-Bench Science 0.1 (Vals) Scores
All 70 tasks from v0.1.0; fixed Terminus 2 harness, pass@1, eight-hour budget. Infrastructure retries permitted. Not interchangeable with native-agent official scores. Standalone raw result; does not contribute to IQ/EQ.

How to read this chart

Bars rank models by the source-backed benchmark value used for this chart. Longer bars indicate higher published scores.

Data sources
Data updated Sep 22, 2026
Terminal-Bench Science 0.1 (Vals) vs Effective Cost
X = effective cost (log). Y = Terminal-Bench Science 0.1 (Vals). All 70 tasks from v0.1.0; fixed Terminus 2 harness, pass@1, eight-hour budget. Infrastructure retries permitted. Not interchangeable with native-agent official scores.
Controls:
Terminal-Bench Science 0.1 (Vals)1:1Cost

How to read this chart

Each point is a public model. The chart compares Terminal-Bench Science 0.1 (Vals) (%) against Effective Cost (per 1M I/O Tokens), with color showing the model provider.