findeverythingaiTHE INDEPENDENT BENCHMARK INDEX

AI, WITH THE RECEIPTS.

Know the score.
See the context.

Explore published model and agent results. Compare the test setup, the measured performance and the reported cost—then read the original source.

Independent index · Coding & terminal tasks · No sponsored rankings

5Separate test suites
222Published benchmark runs
922Research & open-source posts
100% sourcedEvery result links to its origin

01 / THE BENCHMARK DESK

A score needs a setting.

Pick a test suite. Search the runs. Add up to three to your comparison.

21 published runs · Terminal-Bench 4.0

  1. Filtered rank 1

    Opus 5.5

    Claude Code · 2.1.284Source run ↗
    max · 330 trials
    64.8%±3.11 pp
    95% CI half-width
  2. Filtered rank 2

    Sonnet 5.5

    Claude Code · 2.1.284Source run ↗
    max · 330 trials
    61.8%±2.94 pp
    95% CI half-width
  3. Filtered rank 3

    GPT-6.1 Sol

    Codex · 0.159.0Source run ↗
    max · 330 trials
    58.2%±3.11 pp
    95% CI half-width
  4. Filtered rank 4

    Fable 5.1

    Claude Code · 2.1.257Source run ↗
    max · 330 trials
    57.9%±3.76 pp
    95% CI half-width
  5. Filtered rank 5

    Opus 5

    Claude Code · 2.1.231Source run ↗
    max · 330 trials
    51.8%±3.39 pp
    95% CI half-width
  6. Filtered rank 6

    GPT-6 Sol

    Codex · 0.159.0Source run ↗
    max · 330 trials
    49.4%±3.23 pp
    95% CI half-width
  7. Filtered rank 7

    Fable 5

    Claude Code · 2.1.231Source run ↗
    max · 330 trials
    44.5%±3.85 pp
    95% CI half-width
  8. Filtered rank 8

    GLM-5.3

    Claude Code · 2.1.207Source run ↗
    max · 330 trials
    41.8%±3.23 pp
    95% CI half-width
  9. Filtered rank 9

    GPT-5.6 Sol

    Codex · 0.149.1Source run ↗
    max · 330 trials
    37.3%±3.78 pp
    95% CI half-width
  10. Filtered rank 10

    GLM-5.3-Flash

    Claude Code · 2.1.285Source run ↗
    none · 330 trials
    35.8%±3.54 pp
    95% CI half-width
Page 1

Ranks are within the filtered cohort. USD costs are publisher-reported evaluation totals, not API prices. A run’s agent, settings and source date matter. An unlisted model has no imported result, not a score of zero.

02 / SIDE BY SIDE

Your comparison desk.

Same suite, same task cohort. Different agent settings remain visible.

Select runs from the benchmark list.

Start with a real question.

Which setup scored higher? What did the evaluation cost? Add two or three runs to see the details together.

These are published evaluations, not tests conducted by FindEverythingAI. We do not combine unrelated benchmarks into a single “best AI” score. Changes in harness version, prompts and settings can affect results. Overlapping uncertainty ranges do not establish a clear winner.

03 / THE READING ROOM

Research, from the source.

Free, licensed RSS metadata from arXiv and Creative Commons. Original writing stays with its publisher.

A feed you can take with you.

New research announcements and open-source posts, with attribution.

Subscribe via RSS ↗

922 posts · publisher metadata

Feed licences & freshness ↗
Page 1

arXiv items are research announcements and may be preprints, not peer-reviewed findings. We merge duplicate links across topics. arXiv dates are feed announcement dates; Creative Commons dates come from its Atom feed. Benchmarks refresh daily; feeds are checked hourly.

04 / READ THE FINE PRINT

Evidence before the headline.

Keep the tests separate.

Terminal tasks, code editing and refactoring measure different work. Even different versions of the same benchmark are separate comparisons.

Check what was tested.

The model, agent, reasoning effort, task count and evaluation date all affect the score. A stronger result on one suite is not a universal recommendation.

Follow the provenance.

Every run links to a revision of its publisher’s data. We show refresh status, licences and historical coverage, and retain earlier data if a refresh fails.

Read the methodology & source register ↗