findeverythingaiTHE INDEPENDENT BENCHMARK INDEX

BENCHMARK GLOSSARY

The terms behind the table

Published 8 October 2026 · Public product and editorial information

Use these definitions while reading a published run. A publisher may define a term more precisely; its experiment documentation takes precedence.

Benchmark suite
A collection of evaluation tasks and the rules used to score them. Different suites can measure different capabilities.
Task cohort
The particular group or count of tasks evaluated in a run. Equal task counts alone do not prove identical conditions.
Agent or harness
The software that gives a model instructions, tools and an execution environment. The model and the surrounding system both matter.
Run
A recorded evaluation of a particular setup. Several runs may use the same model with different settings.
Score
The reported outcome under the suite’s metric. On this desk, percentage scores are shown with the publisher’s setup and source.
Reasoning effort
A reported model setting that can change the computation used for a response. Labels and behavior are provider-specific.
Reported spend
The publisher’s evaluation cost where supplied. It is distinct from an advertised API rate or your future monthly bill.
Uncertainty
Reported variation or a statistical interval. Different kinds of intervals are not interchangeable, and missing values are not invented.
Historical result
An older evaluation retained for context. A recent refresh does not turn it into a recent test.
Missing result
An unavailable record or field. Missing is not zero and is not evidence that a model failed.
Source check
The time the importer checked a publisher. It is not the time of the underlying experiment.
Stale source
A source without a sufficiently recent successful refresh under the importer’s policy. Retained records can remain readable while freshness is uncertain.

Read a result step by step →Inspect sources and refresh rules →