Long Context Reasoning (AA-LCR) Extended

GPT-5.6 Sol
GPT-5.6 SolMax • Pro
79.5%
GPT-5.6 Terra
GPT-5.6 TerraMax
74.4%
Opus 5.5
Opus 5.5Max
74.2%
Opus 5
Opus 5Max
73.9%
Sonnet 5.5
Sonnet 5.5Max
73.8%
Fable 5.1
Fable 5.1High
73.0%
GPT-5.6 Luna
GPT-5.6 LunaMax
72.2%
Fable 5
Fable 5Max
71.1%
DeepSeek-V4.1-Flash
DeepSeek-V4.1-FlashMax
70.8%
GPT-5.5
GPT-5.5xHigh
70.5%
GPT-6 Sol
GPT-6 SolMax
70.2%
Kimi K3
Kimi K3Max
69.1%
GPT-6 Luna
GPT-6 LunaMax
68.3%
GPT-5.4
GPT-5.4xHigh
68.3%
Muse Spark 1.3
Muse Spark 1.3Max
66.9%
Grok 4.6
Grok 4.6xHigh
65.2%
Opus 4.8
Opus 4.8Max
63.2%
Sonnet 5
Sonnet 5Max
61.8%
DeepSeek-V4-Flash
DeepSeek-V4-FlashMax
61.8%
Muse Spark 1.2
Muse Spark 1.2xHigh
61.5%
Grok 4.5
Grok 4.5High
60.4%
Opus 4.7
Opus 4.7Max
59.8%
Sonnet 4.6
Sonnet 4.6High
58.1%
Inkling
InklingHigh
50.6%
Kimi K2.7 Code
Kimi K2.7 CodeHigh
49.2%
Gemini 3.6 Flash
Gemini 3.6 FlashHigh
46.3%
MiniMax-M3
MiniMax-M3High
45.5%
GLM-5.2
GLM-5.2Max
42.7%
Nemotron 3 Ultra
Nemotron 3 UltraHigh
41.6%
Gemini 3.5 Flash
Gemini 3.5 FlashHigh
39.9%
DeepSeek-V4-Pro
DeepSeek-V4-ProMax
37.6%
Gemini 3.1 Pro
Gemini 3.1 ProHigh
35.4%
Qwen3.5
Qwen3.5
33.4%
Kimi K2
Kimi K2High • Thinking
32.9%
MiniMax-M2.7
MiniMax-M2.7High
28.4%
GPT-OSS-120B
GPT-OSS-120BHigh
25.8%
DeepSeek-V3.2
DeepSeek-V3.2
24.7%
Gemma 4 31B
Gemma 4 31B
22.7%
10%
20%
30%
40%
50%
60%
70%
80%
90%
100%

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.