DeepSWE

Measuring frontier coding agents on original, long-horizon engineering tasks

Get notified when new models drop

Leaderboard

113 tasks · updated August 20, 2026 · changelog
DeepSWE score$0$5.00$100%10%20%30%40%50%60%70%80%Avg cost per taskmost efficient ↗gpt-5.6-lunaMEDIUMgemini-3.5-flashHIGHglm-5.2MAXgemini-3.6-flashHIGHclaude-sonnet-5HIGHclaude-opus-4.8HIGHdeepseek-v4-flashMAXgpt-5.5MEDIUMmuse-spark-1.2XHIGHqwen3.8-maxXHIGHgpt-5.6-solMEDIUMdeepseek-v4-proMAXgemini-3.7-flashMEDIUMgrok-4.6XHIGHkimi-k3MAXclaude-fable-5HIGHglm-5.3MAXclaude-opus-5MAX
claude-opus-5[max]
74%Âą4%
Avg cost $11.84Out tok 118kSteps 99
gpt-5.6-sol[max]
73%Âą3%
Avg cost $6.46Out tok 60kSteps 61
claude-fable-5[max]
70%Âą4%
Avg cost $21.63Out tok 119kSteps 88
glm-5.3[max]
69%Âą3%
Avg cost $3.99Out tok 80kSteps 124
kimi-k3[max]
69%Âą5%
Avg cost $4.65Out tok 81kSteps 98
gpt-5.6-luna[max]
67%Âą4%
Avg cost $0.61Out tok 73kSteps 102
gpt-5.5[xhigh]
67%Âą6%
Avg cost $7.23Out tok 46kSteps 82
grok-4.6[xhigh]
67%Âą2%
Avg cost $5.50Out tok 71kSteps 87
gemini-3.7-flash[high]
65%Âą2%
Avg cost $2.18Out tok 107kSteps 125
deepseek-v4-pro[max]
63%Âą6%
Avg cost $1.67Out tok 106kSteps 155
claude-opus-4.8[max]
59%Âą2%
Avg cost $13.22Out tok 135kSteps 120
qwen3.8-max[xhigh]
57%Âą3%
Avg cost $3.73Out tok 95kSteps 111
muse-spark-1.2[xhigh]
55%Âą2%
Avg cost $3.70Out tok 99kSteps 101
claude-sonnet-5[max]
54%Âą4%
Avg cost $26.40Out tok 214kSteps 268
deepseek-v4-flash[max]
53%Âą4%
Avg cost $0.46Out tok 108kSteps 153
gemini-3.6-flash[high]
47%Âą4%
Avg cost $2.21Out tok 96kSteps 117
glm-5.2[max]
44%Âą2%
Avg cost $3.92Out tok 78kSteps 129
gemini-3.5-flash[high]
36%Âą4%
Avg cost $3.45Out tok 76kSteps 105
0%20%40%60%80%

All models run on mini-swe-agent for consistency. Read why.

Today's leading public coding benchmarks are starting to saturate at the frontier: top models cluster within a narrow score band where adjacent configurations often overlap on confidence intervals. DeepSWE is a long-horizon software engineering benchmark built to separate them. It delivers four advances over existing public benchmarks:

  • Contamination free: Tasks are written from scratch, not adapted from existing commits or PRs, so no model has seen the solution during pretraining.
  • High diversity: Tasks span a broad pool of 91 repositories across 5 languages.
  • Real-world complexity: Prompts are ~half the length of SWE-bench Pro's, yet solutions require 5.5x more code and ~2x more output tokens.
  • Reliable verification: Verifiers are hand-written to test software behavior rather than implementation details.

The result is a benchmark that reflects how today's frontier coding agents actually perform in software engineering work.

Task Examples

All 113 tasks

Sign up for leaderboard updates

New frontier models are added to the DeepSWE leaderboard as they're released.