PUBLISHED BENCHMARK RESULTS
Terminal-Bench 4.0
Agent and model runs on terminal tasks. Published accuracy, with 95% confidence intervals where reported.
These are publisher-reported model and agent runs. Agent version, reasoning effort or editing format, environment and source dates may differ. Task counts do not prove identical evaluation conditions. Compare setups before drawing conclusions.
Reported costs are evaluation totals in USD, not current API prices. Missing results are not scores of zero. No score is a universal recommendation.
Benchmark publisher ↗ · Apache-2.0 ↗ · Full methodology and provenance
Published suite
- Filtered rank 1
Opus 5.5
Source run ↗max64.8%±3.11 ppCompare
95% CI half-width - Filtered rank 2
Sonnet 5.5
Source run ↗max61.8%±2.94 ppCompare
95% CI half-width - Filtered rank 3
GPT-6.1 Sol
Source run ↗max58.2%±3.11 ppCompare
95% CI half-width - Filtered rank 4
Fable 5.1
Source run ↗max57.9%±3.76 ppCompare
95% CI half-width - Filtered rank 5
Opus 5
Source run ↗max51.8%±3.39 ppCompare
95% CI half-width - Filtered rank 6
GPT-6 Sol
Source run ↗max49.4%±3.23 ppCompare
95% CI half-width - Filtered rank 7
Fable 5
Source run ↗max44.5%±3.85 ppCompare
95% CI half-width - Filtered rank 8
GLM-5.3
Source run ↗max41.8%±3.23 ppCompare
95% CI half-width - Filtered rank 9
GPT-5.6 Sol
Source run ↗max37.3%±3.78 ppCompare
95% CI half-width - Filtered rank 10
GLM-5.3-Flash
Source run ↗none35.8%±3.54 ppCompare
95% CI half-width - Filtered rank 11
Qwen3.8-Max-0902
Source run ↗max27.0%±3.78 ppCompare
95% CI half-width - Filtered rank 12
Opus 4.8
Source run ↗max23.6%±3.56 ppCompare
95% CI half-width - Filtered rank 13
GPT-5.6 Terra
Source run ↗max21.5%±3.25 ppCompare
95% CI half-width - Filtered rank 14
Grok 4.6
Source run ↗none20.3%±3.09 ppCompare
95% CI half-width - Filtered rank 15
Gemini 3.8 Flash
Source run ↗high19.1%±3.36 ppCompare
95% CI half-width - Filtered rank 16
GPT-5.6 Luna
Source run ↗max17.3%±2.85 ppCompare
95% CI half-width - Filtered rank 17
GPT-6 Luna
Source run ↗max16.4%±2.72 ppCompare
95% CI half-width - Filtered rank 18
Muse Spark 1.3
Source run ↗xhigh14.6%±2.94 ppCompare
95% CI half-width - Filtered rank 19
Grok 4.5
Source run ↗none12.4%±2.62 ppCompare
95% CI half-width - Filtered rank 20
Sonnet 5
Source run ↗max12.4%±3.06 ppCompare
95% CI half-width - Filtered rank 21
Gemini 3.7 Flash
Source run ↗high11.2%±2.45 ppCompare
95% CI half-width