SWE-bench Verified

Fable 5
Fable 5Max
95.9%
Opus 5
Opus 5Max
93.5%
Fable 5.1
Fable 5.1High
92.1%
Opus 4.8
Opus 4.8Max
88.5%
Opus 4.7
Opus 4.7Max
83.1%
DeepSeek-V4.1-Flash
DeepSeek-V4.1-FlashMax
82.9%
Sonnet 5
Sonnet 5Max
82.8%
Grok 4.5
Grok 4.5High
81.7%
GLM-5.3
GLM-5.3Max
81.7%
Gemini 3.6 Flash
Gemini 3.6 FlashHigh
81.1%
Gemini 3.7 Flash
Gemini 3.7 FlashHigh
81.0%
Gemini 3.8 Flash
Gemini 3.8 FlashHigh
80.6%
Sonnet 4.6
Sonnet 4.6High
79.1%
GPT-5.5
GPT-5.5xHigh
78.8%
Gemini 3.5 Flash
Gemini 3.5 FlashHigh
78.5%
GLM-5.2
GLM-5.2Max
78.5%
Gemini 3.1 Pro
Gemini 3.1 ProHigh
78.4%
Kimi K2.7 Code
Kimi K2.7 CodeHigh
78.2%
GPT-5.4
GPT-5.4xHigh
78.1%
MiniMax-M3
MiniMax-M3High
75.5%
DeepSeek-V4-Pro
DeepSeek-V4-ProMax
75.4%
MiniMax-M2.7
MiniMax-M2.7High
75.3%
Qwen3.5
Qwen3.5
72.1%
DeepSeek-V3.2
DeepSeek-V3.2
66.0%
Gemma 4 31B
Gemma 4 31B
46.1%
40%
50%
60%
70%
80%
90%
100%

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.