findeverythingai← Benchmark desk

THE SOURCE REGISTER

Show the work.

FindEverythingAI is an independent index of published benchmark runs and licensed research metadata. We import results from the original repositories; we do not run the evaluations ourselves. Our initial coverage is coding and terminal agents. It is not a complete assessment of reasoning, image generation, general knowledge or every AI model.

What a comparison means

Benchmark sources & freshness

Benchmark snapshots are checked daily. Files are read from a pinned Git commit for each import, and each run links to that revision. The snapshot checksum covers the imported result files. Normalization selects display fields and converts missing values; it does not revise the publisher’s scores.

Loading source status…

Original repositories: Terminal-Bench, Terminal-Bench 2.1 and Aider. See Aider’s benchmark documentation. Benchmark imports are attributed to their publishers under Apache-2.0, with original licence copies available above. Our transformation is an independent presentation of those datasets. Project names and trademarks do not imply affiliation.

RSS reuse policy

Being available through RSS does not by itself permit republication. We only display feed metadata from sources with explicit reuse permission. We do not copy full papers, abstracts, blog articles or publisher images, and we do not load tracking images from feeds.

Feed sources & freshness

Feeds are checked hourly using conditional requests where supported. Duplicate article links across arXiv categories are merged, with their topics retained. Up to 35 days of research announcements are retained, with a cap of 2,000 posts across sources. Creative Commons’ published feed archive is also retained. Our RSS feed exports the latest 100 entries with attribution.

Loading feed status…

Original feeds: arXiv AI, arXiv language, arXiv machine learning and Creative Commons Open Source.

When a source fails

A failed refresh retains the last successful snapshot and marks that source stale. A source without any successful import is unavailable. The browser also labels benchmark data stale after 27 hours without a successful check, and feed data after three hours. Source failures are not reported as empty successful evaluations. Publisher corrections appear on the next successful import.

Downloads

The normalized benchmark snapshot and feed metadata snapshot are available as JSON. CSV export includes only the runs matching your filters, along with original result links. Preserve the applicable source licences and attribution when reusing data.

Back to the benchmark desk