findeverythingaiInteractive benchmark desk

PUBLISHED BENCHMARK RESULTS

Terminal-Bench 2.1

A separate terminal task suite. Published accuracy; uncertainty is standard error, not a 95% confidence interval.

These are publisher-reported model and agent runs. Agent version, reasoning effort or editing format, environment and source dates may differ. Task counts do not prove identical evaluation conditions. Compare setups before drawing conclusions.

Reported costs are evaluation totals in USD, not current API prices. Missing results are not scores of zero. No score is a universal recommendation.

Latest source result: 2026-07-11 · Source check: 2026-10-08 · Source checked successfully

Benchmark publisher ↗ · Apache-2.0 ↗ · Full methodology and provenance

Published suite

20 runs · Task success. Ranks apply within this task cohort.

  1. Filtered rank 1

    Fable 5

    Claude Code · 2.1.167Source run ↗
    xhigh · 445 trials
    83.8%±1.16 pp
    Standard error
    Compare
  2. Filtered rank 2

    GPT-5.5

    Codex · 0.125.0Source run ↗
    xhigh · 445 trials
    83.2%±1.13 pp
    Standard error
    Compare
  3. Filtered rank 3

    Fable 5

    Terminus 2 · 2.0.0Source run ↗
    high · 445 trials
    80.5%±1.16 pp
    Standard error
    Compare
  4. Filtered rank 4

    Grok 4.5

    Cursor CLI · 2026.07.08-0c04a8aSource run ↗
    high · 445 trials
    79.3%±1.46 pp
    Standard error
    Compare
  5. Filtered rank 5

    Opus 4.8

    Claude Code · 2.1.205Source run ↗
    high · 445 trials
    78.9%±1.31 pp
    Standard error
    Compare
  6. Filtered rank 6

    GPT-5.6 Terra

    Codex · 0.144.1Source run ↗
    max · 445 trials
    78.4%±1.25 pp
    Standard error
    Compare
  7. Filtered rank 7

    GPT-5.5

    Terminus 2 · 2.0.0Source run ↗
    xhigh · 445 trials
    78.0%±1.22 pp
    Standard error
    Compare
  8. Filtered rank 8

    GPT-5.6 Sol

    Codex · 0.144.0Source run ↗
    max · 445 trials
    76.2%±1.28 pp
    Standard error
    Compare
  9. Filtered rank 9

    Muse Spark 1.1

    mini-SWE-agent · 2.4.5Source run ↗
    xhigh · 445 trials
    76.2%±1.23 pp
    Standard error
    Compare
  10. Filtered rank 10

    GPT-5.6 Luna

    Codex · 0.144.1Source run ↗
    max · 445 trials
    75.7%±1.32 pp
    Standard error
    Compare
  11. Filtered rank 11

    Sonnet 5

    Claude Code · 2.1.205Source run ↗
    high · 445 trials
    74.6%±1.64 pp
    Standard error
    Compare
  12. Filtered rank 12

    GPT-5.6 Terra

    Codex · 0.144.0Source run ↗
    max · 445 trials
    74.4%±1.54 pp
    Standard error
    Compare
  13. Filtered rank 13

    Gemini 3 Pro

    Terminus 2 · 2.0.0Source run ↗
    high · 445 trials
    73.9%±1.29 pp
    Standard error
    Compare
  14. Filtered rank 14

    GPT-5.6 Luna

    Codex · 0.144.0Source run ↗
    max · 445 trials
    71.2%±1.39 pp
    Standard error
    Compare
  15. Filtered rank 15

    Opus 4.7

    Claude Code · 2.1.123Source run ↗
    max · 447 trials
    68.9%±1.41 pp
    Standard error
    Compare
  16. Filtered rank 16

    Opus 4.7

    Terminus 2 · 2.0.0Source run ↗
    max · 445 trials
    66.1%±1.37 pp
    Standard error
    Compare
  17. Filtered rank 17

    Gemini 3 Pro

    Gemini CLI · 0.40.0Source run ↗
    high · 445 trials
    65.8%±1.38 pp
    Standard error
    Compare
  18. Filtered rank 18

    Gemini 3.1 Pro

    Gemini CLI · 0.40.0Source run ↗
    high · 445 trials
    65.8%±1.67 pp
    Standard error
    Compare
  19. Filtered rank 19

    Gemini 3.1 Pro

    Terminus 2 · 2.0.0Source run ↗
    high · 445 trials
    65.6%±1.65 pp
    Standard error
    Compare
  20. Filtered rank 20

    GLM-5.1

    Claude Code · 2.1.123Source run ↗
    max · 445 trials
    58.6%±1.24 pp
    Standard error
    Compare