PUBLISHED BENCHMARK RESULTS
Aider Refactoring
Historical refactoring runs using Aider. Scores use one attempt, unlike the other Aider suites.
These are publisher-reported model and agent runs. Agent version, reasoning effort or editing format, environment and source dates may differ. Task counts do not prove identical evaluation conditions. Compare setups before drawing conclusions.
Reported costs are evaluation totals in USD, not current API prices. Missing results are not scores of zero. No score is a universal recommendation.
Historical coverage: this source contains older Aider evaluations. Refreshing the source does not make the underlying results current.
Benchmark publisher ↗ · Apache-2.0 ↗ · Full methodology and provenance
89 tasks
- Filtered rank 1
claude-3-5-sonnet-20241022
Source run ↗diff92.1%Compare - Filtered rank 2
o1-preview
Source run ↗diff75.3%Compare - Filtered rank 3
claude-3.5-sonnet-20240620
Source run ↗diff64.0%Compare - Filtered rank 4
gpt-4o
Source run ↗diff62.9%Compare - Filtered rank 5
gpt-4-1106-preview
Source run ↗udiff50.6%Compare - Filtered rank 6
gemini/gemini-1.5-pro-latest
Source run ↗diff-fenced49.4%Compare - Filtered rank 7
gpt-4o-2024-08-06
Source run ↗diff49.4%Compare - Filtered rank 8
o1-mini
Source run ↗diff44.9%Compare - Filtered rank 9
gpt-4-0125-preview
Source run ↗udiff33.7%Compare - Filtered rank 10
DeepSeek Coder V2 0724 (deprecated)
Source run ↗diff32.6%Compare - Filtered rank 11
DeepSeek Chat V2.5
Source run ↗diff31.5%Compare
88 tasks
- Filtered rank 1
gpt-4-turbo-2024-04-09 (udiff)
Source run ↗udiff34.1%Compare - Filtered rank 2
gpt-4-turbo-2024-04-09 (diff)
Source run ↗diff21.4%Compare
83 tasks
- Filtered rank 1
claude-3-opus-20240229
Source run ↗diff72.3%Compare