BENCHMARK GLOSSARY
The terms behind the table
Use these definitions while reading a published run. A publisher may define a term more precisely; its experiment documentation takes precedence.
- Benchmark suite
- A collection of evaluation tasks and the rules used to score them. Different suites can measure different capabilities.
- Task cohort
- The particular group or count of tasks evaluated in a run. Equal task counts alone do not prove identical conditions.
- Agent or harness
- The software that gives a model instructions, tools and an execution environment. The model and the surrounding system both matter.
- Run
- A recorded evaluation of a particular setup. Several runs may use the same model with different settings.
- Score
- The reported outcome under the suite’s metric. On this desk, percentage scores are shown with the publisher’s setup and source.
- Reasoning effort
- A reported model setting that can change the computation used for a response. Labels and behavior are provider-specific.
- Reported spend
- The publisher’s evaluation cost where supplied. It is distinct from an advertised API rate or your future monthly bill.
- Uncertainty
- Reported variation or a statistical interval. Different kinds of intervals are not interchangeable, and missing values are not invented.
- Historical result
- An older evaluation retained for context. A recent refresh does not turn it into a recent test.
- Missing result
- An unavailable record or field. Missing is not zero and is not evidence that a model failed.
- Source check
- The time the importer checked a publisher. It is not the time of the underlying experiment.
- Stale source
- A source without a sufficiently recent successful refresh under the importer’s policy. Retained records can remain readable while freshness is uncertain.
Read a result step by step →Inspect sources and refresh rules →