THE SOURCE REGISTER
Show the work.
FindEverythingAI is an independent index of published benchmark runs and licensed research metadata. We import results from the original repositories; we do not run the evaluations ourselves. Our initial coverage is coding and terminal agents. It is not a complete assessment of reasoning, image generation, general knowledge or every AI model.
What a comparison means
- One suite at a time. Terminal-Bench 4.0 and 2.1 are different task suites. The three Aider benchmarks are separate. Scores are never averaged across suites.
- One task cohort at a time. Aider datasets contain runs with different task counts. Selection and sharing require the same imported benchmark and reported task count. Equal task counts alone do not prove identical prompts, environments or task membership.
- Models plus agents. Terminal-Bench scores describe a model running through a particular agent. Aider scores describe its coding harness and editing format. We preserve the reported harness version and settings where available. Differences in these settings can confound a model comparison.
- Published scores. We preserve the publisher’s accuracy or pass rate as a percentage. Terminal-Bench 4.0 publishes a 95% confidence interval half-width; 2.1 publishes standard error. These are different uncertainty measures, both shown in percentage points. Missing uncertainty is labelled “not reported.”
- Attempts matter. Aider Polyglot and Code Editing use the reported pass rate after two attempts. Aider Refactoring uses the first-attempt pass rate. Terminal-Bench trial counts can include repeated runs; they are not unique task counts.
- Cost is evaluation spend. Reported cost is the publisher’s total USD evaluation cost, not a subscription or API price. A derived cost per trial is total spend divided by reported trials. Aider’s zero-cost values are treated as unreported, following its leaderboard display. Costs across different trials and environments need context.
- Dates are source dates. We keep the source metadata date, which may be a release or submission date. It is not necessarily the day every task ran. “Last successful check” describes our import time. Historical Aider coverage remains dated as historical even when refreshed today.
- Missing is not zero. An unlisted model has no result in our imported sources. A higher benchmark score does not guarantee it will perform better on your own work.
Benchmark sources & freshness
Benchmark snapshots are checked daily. Files are read from a pinned Git commit for each import, and each run links to that revision. The snapshot checksum covers the imported result files. Normalization selects display fields and converts missing values; it does not revise the publisher’s scores.
Loading source status…
Original repositories: Terminal-Bench, Terminal-Bench 2.1 and Aider. See Aider’s benchmark documentation. Benchmark imports are attributed to their publishers under Apache-2.0, with original licence copies available above. Our transformation is an independent presentation of those datasets. Project names and trademarks do not imply affiliation.
RSS reuse policy
Being available through RSS does not by itself permit republication. We only display feed metadata from sources with explicit reuse permission. We do not copy full papers, abstracts, blog articles or publisher images, and we do not load tracking images from feeds.
- arXiv: titles, authors, announcement dates and original links from cs.AI, cs.CL and cs.LG. The arXiv metadata policy makes metadata available under CC0. Individual papers can have different licences; CC0 metadata does not grant reuse of the paper itself. arXiv items may be unreviewed preprints.
- Creative Commons Open Source: titles, author names, feed dates and original links. The publisher’s site is licensed under CC BY 4.0 except where otherwise noted. Each entry includes publisher attribution, the original URL and licence link. We omit article content and normalize metadata presentation.
Feed sources & freshness
Feeds are checked hourly using conditional requests where supported. Duplicate article links across arXiv categories are merged, with their topics retained. Up to 35 days of research announcements are retained, with a cap of 2,000 posts across sources. Creative Commons’ published feed archive is also retained. Our RSS feed exports the latest 100 entries with attribution.
Loading feed status…
Original feeds: arXiv AI, arXiv language, arXiv machine learning and Creative Commons Open Source.
When a source fails
A failed refresh retains the last successful snapshot and marks that source stale. A source without any successful import is unavailable. The browser also labels benchmark data stale after 27 hours without a successful check, and feed data after three hours. Source failures are not reported as empty successful evaluations. Publisher corrections appear on the next successful import.
Downloads
The normalized benchmark snapshot and feed metadata snapshot are available as JSON. CSV export includes only the runs matching your filters, along with original result links. Preserve the applicable source licences and attribution when reusing data.
Back to the benchmark desk