Long Context Reasoning (AA-LCR)

GPT-5.5
GPT-5.5xHigh
82.8%
GPT-5.4
GPT-5.4xHigh
82.5%
GPT-5.6 Terra
GPT-5.6 TerraMax
81.8%
Fable 5
Fable 5Max
81.5%
Kimi K3
Kimi K3Max
81.5%
Fable 5.1
Fable 5.1High
81.5%
Gemini 3.6 Flash
Gemini 3.6 FlashHigh
81.2%
DeepSeek-V4.1-Flash
DeepSeek-V4.1-FlashMax
81.0%
GPT-5.6 Luna
GPT-5.6 LunaMax
80.7%
Opus 5
Opus 5Max
80.5%
Kimi K2.7 Code
Kimi K2.7 CodeHigh
80.5%
GPT-5.6 Sol
GPT-5.6 SolMaxPro
80.0%
MiniMax-M3
MiniMax-M3High
80.0%
Gemini 3.1 Pro
Gemini 3.1 ProHigh
79.2%
GLM-5.2
GLM-5.2Max
78.7%
GPT-6 Astra
GPT-6 AstraxHigh
78.2%
DeepSeek-V4-Flash
DeepSeek-V4-FlashMax
78.0%
Qwen3.5
Qwen3.5
78.0%
Grok 4.5
Grok 4.5High
77.5%
MiniMax-M2.7
MiniMax-M2.7High
77.5%
Opus 4.7
Opus 4.7Max
77.0%
Sonnet 5
Sonnet 5Max
76.7%
Sonnet 4.6
Sonnet 4.6High
76.5%
Opus 4.8
Opus 4.8Max
76.5%
DeepSeek-V4-Pro
DeepSeek-V4-ProMax
76.2%
Inkling
InklingHigh
75.5%
Gemma 4 31B
Gemma 4 31B
73.3%
DeepSeek-V3.2
DeepSeek-V3.2
70.3%
Kimi K2
Kimi K2HighThinking
67.3%
Nemotron 3 Ultra
Nemotron 3 UltraHigh
67.3%
Gemini 3.5 Flash
Gemini 3.5 FlashHigh
67.0%
Muse Spark 1.1
Muse Spark 1.1xHigh
67.0%
GPT-OSS-120B
GPT-OSS-120BHigh
54.2%
40%
50%
60%
70%
80%
90%
100%

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.