DeepResearch Bench II

Opus 5
Opus 5High
56.1%
Fable 5.1
Fable 5.1Max
55.0%
Sonnet 4.6
Sonnet 4.6High
53.3%
Opus 5.5
Opus 5.5Max
53.3%
DeepSeek V4.1 Flash
DeepSeek V4.1 FlashMax
53.2%
GLM 5.3
GLM 5.3Max
52.0%
Sonnet 5.5
Sonnet 5.5Max
51.7%
GPT 5.6 Sol
GPT 5.6 SolMedium
51.5%
Qwen 3.8 Max
Qwen 3.8 MaxxHigh
51.0%
GPT 6 Astra
GPT 6 AstraxHigh
50.1%
GPT 5.6 Luna
GPT 5.6 LunaMax
49.7%
Grok 4.7
Grok 4.7xHigh
49.0%
GPT 5.6 Terra
GPT 5.6 TerraMedium
47.2%
Grok 4.5
Grok 4.5High
47.1%
Opus 4.7
Opus 4.7High
45.4%
GPT 5.5
GPT 5.5Medium
45.4%
Opus 4.8
Opus 4.8High
44.6%
DeepSeek V4 Pro
DeepSeek V4 ProHigh
44.2%
Gemini 3.7 Flash
Gemini 3.7 FlashMedium
43.4%
GPT 6 Luna
GPT 6 LunaMax
42.9%
GPT 6 Sol
GPT 6 SolMax
42.9%
Sonnet 5
Sonnet 5High
39.7%
GPT 5.4
GPT 5.4None
38.4%
Gemini 3.6 Flash
Gemini 3.6 FlashMedium
36.3%
30%
40%
50%
60%
70%
80%

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.