ABOUT THE COMPARISON DESK
A score is the beginning of a question
FindEverythingAI is an independent reading and comparison desk for published AI evaluations. It helps you inspect a result in context: the task set, agent, setup, date, source and reported cost.
What is measured—and by whom
The benchmark publishers run the evaluations shown here. FindEverythingAI normalizes their published result records; it does not claim to have independently repeated each run. The methodology page identifies sources, licences and refresh behavior.
Compare like with like
A score is meaningful within its suite and task cohort. Agent version, reasoning effort, environment and evaluation cost can differ even when the model label looks similar. The comparison controls preserve suite and cohort boundaries; the reading guide explains the remaining caveats.
Publication and commercial policy
Imported research entries are attributed metadata linked to their original publishers. Inclusion is not an endorsement. No paid placement is currently part of benchmark ranking. If commercial placements are introduced, they must be labelled and separate from the reported results.
Report a discrepancy
Use the website enquiry desk, select “Discuss” and include FindEverythingAI, the page URL, the source-run URL and the field that appears wrong. Do not send private data. Corrections should be checked against the publisher’s evidence.