Terminal-Bench 2.1 Extended

DeepSeek-V4.1-Flash
DeepSeek-V4.1-FlashMax
38.4%
GPT-5.5
GPT-5.5xHigh
38.0%
GPT-5.6 Sol
GPT-5.6 SolxHigh
37.4%
Gemini 3.1 Pro
Gemini 3.1 ProHigh
35.0%
Opus 5
Opus 5High
34.0%
Grok 4.5
Grok 4.5High
33.3%
GPT-5.4
GPT-5.4xHigh
31.3%
Kimi K3
Kimi K3Max
31.0%
DeepSeek-V4-Flash
DeepSeek-V4-FlashMax
30.0%
Opus 4.8
Opus 4.8High
29.3%
Sonnet 5
Sonnet 5High
28.9%
Gemini 3.5 Flash
Gemini 3.5 FlashHigh
28.6%
Opus 4.7
Opus 4.7High
27.6%
Kimi K2.7 Code
Kimi K2.7 CodeHigh
23.6%
MiniMax-M3
MiniMax-M3High
22.2%
DeepSeek-V4-Pro
DeepSeek-V4-ProMax
21.2%
Qwen3.5
Qwen3.5
17.8%
Sonnet 4.6
Sonnet 4.6High
17.5%
Fable 5
Fable 5Max
16.8%
DeepSeek-V3.2
DeepSeek-V3.2
15.2%
Nemotron 3 Ultra
Nemotron 3 UltraHigh
11.8%
Kimi K2
Kimi K2HighThinking
11.1%
MiniMax-M2.7
MiniMax-M2.7High
10.1%
0%
10%
20%
30%
40%
50%
60%
70%
80%
90%
100%

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.