DeepSWE v1.1 — performance / price by model & effort
Source data: https://deepswe.datacurve.ai/artifacts/v1.1/leaderboard-live.json
Accessed July 15, 2026 · Performance = pass@1 rate · Price = avg $ per task · Perf/Price = pass@1 points per $
| Model | Effort | Perf/Price | Performance | Price | Steps |
|---|---|---|---|---|---|
| gpt-5-6-terra | medium | 60.20 | 35.1% | $0.58 | 24 |
| gpt-5-6-luna | high | 56.88 | 44.2% | $0.78 | 44 |
| gpt-5-6-terra | low | 56.23 | 24.1% | $0.43 | 20 |
| gpt-5-6-luna | medium | 52.16 | 11.3% | $0.22 | 22 |
| gpt-5-6-terra | high | 47.39 | 53.8% | $1.13 | 31 |
| gpt-5-6-sol | low | 42.22 | 45.4% | $1.07 | 21 |
| gpt-5-6-luna | xhigh | 37.03 | 56.9% | $1.54 | 63 |
| gpt-5-6-sol | medium | 32.79 | 61.1% | $1.86 | 26 |
| gpt-5-6-terra | xhigh | 28.29 | 60.2% | $2.13 | 39 |
| muse-spark-1-1 | xhigh | 22.58 | 53.3% | $2.36 | 86 |
| gpt-5-5 | low | 22.49 | 27.0% | $1.20 | 27 |
| gpt-5-6-luna | max | 22.19 | 67.2% | $3.03 | 92 |
| gpt-5-6-luna | low | 21.39 | 1.5% | $0.07 | 12 |
| gpt-5-6-sol | high | 20.00 | 69.4% | $3.47 | 32 |
| gpt-5-5 | medium | 19.63 | 54.0% | $2.75 | 43 |
| claude-opus-4-8 | low | 17.79 | 40.8% | $2.29 | 47 |
| claude-fable-5 | low | 15.86 | 59.6% | $3.76 | 29 |
| gpt-5-6-sol | xhigh | 15.04 | 70.7% | $4.70 | 39 |
| claude-opus-4-8 | medium | 14.13 | 48.7% | $3.44 | 60 |
| gpt-5-6-terra | max | 14.08 | 69.6% | $4.95 | 71 |
| claude-sonnet-5 | low | 13.95 | 30.5% | $2.19 | 70 |
| glm-5-2 | high | 12.80 | 36.3% | $2.84 | 112 |
| gpt-5-5 | high | 12.62 | 64.4% | $5.10 | 60 |
| claude-opus-4-8 | high | 12.09 | 51.8% | $4.28 | 67 |
| glm-5-2 | max | 11.17 | 43.8% | $3.92 | 123 |
| kimi-k2-7-code | None | 10.84 | 30.5% | $2.82 | 139 |
| claude-fable-5 | medium | 10.74 | 65.4% | $6.09 | 41 |
| claude-sonnet-5 | medium | 9.75 | 39.8% | $4.08 | 100 |
| gpt-5-5 | xhigh | 9.28 | 67.0% | $7.23 | 76 |
| gpt-5-4 | xhigh | 9.16 | 51.8% | $5.65 | 63 |
| gpt-5-6-sol | max | 8.66 | 72.7% | $8.39 | 53 |
| claude-fable-5 | high | 7.48 | 68.6% | $9.18 | 49 |
| claude-opus-4-8 | xhigh | 6.79 | 54.4% | $8.01 | 90 |
| claude-sonnet-5 | high | 6.50 | 48.2% | $7.43 | 138 |
| claude-sonnet-4-6 | high | 5.42 | 29.9% | $5.52 | 124 |
| claude-fable-5 | xhigh | 5.21 | 69.9% | $13.41 | 61 |
| gemini-3-5-flash | medium | 5.09 | 37.4% | $7.34 | 82 |
| claude-opus-4-8 | max | 4.46 | 59.0% | $13.22 | 116 |
| claude-sonnet-5 | xhigh | 4.18 | 49.7% | $11.89 | 174 |
| claude-fable-5 | max | 3.22 | 69.7% | $21.63 | 79 |
| claude-sonnet-5 | max | 2.04 | 53.8% | $26.40 | 260 |
| gemini-3-1-pro-preview | high | 1.24 | 11.8% | $9.48 | 77 |
42 model/effort configurations from the DeepSWE v1.1 live leaderboard, ranked by performance-per-dollar.