SWE-bench Verified Extended

DeepSeek-V4.1-Flash
DeepSeek-V4.1-FlashMax
82.7%
Opus 5
Opus 5Max
82.0%
Fable 5
Fable 5Max
78.7%
Grok 4.5
Grok 4.5High
70.7%
GLM-5.3
GLM-5.3Max
68.0%
Sonnet 5
Sonnet 5Max
67.3%
Opus 4.8
Opus 4.8Max
63.3%
Opus 4.7
Opus 4.7Max
58.0%
Fable 5.1
Fable 5.1High
57.3%
Gemini 3.8 Flash
Gemini 3.8 FlashHigh
54.0%
GLM-5.2
GLM-5.2Max
47.3%
GPT-5.4
GPT-5.4xHigh
40.0%
GPT-5.5
GPT-5.5xHigh
39.3%
Gemini 3.6 Flash
Gemini 3.6 FlashHigh
33.3%
Gemini 3.7 Flash
Gemini 3.7 FlashHigh
30.0%
MiniMax-M3
MiniMax-M3High
28.7%
Sonnet 4.6
Sonnet 4.6High
27.3%
Kimi K2.7 Code
Kimi K2.7 CodeHigh
26.7%
Gemini 3.5 Flash
Gemini 3.5 FlashHigh
22.7%
DeepSeek-V4-Pro
DeepSeek-V4-ProMax
22.0%
Gemini 3.1 Pro
Gemini 3.1 ProHigh
21.3%
DeepSeek-V3.2
DeepSeek-V3.2
10.0%
MiniMax-M2.7
MiniMax-M2.7High
6.7%
0%
10%
20%
30%
40%
50%
60%
70%
80%
90%
100%

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.