Real-world code issue resolution.
Tests whether LLMs can resolve real-world GitHub issues by editing codebases, spanning 500 problems from 12 popular Python repos and requiring multi-file, long-context changes.
Model
Score
DeepSeek-V4.1-FlashMax
Score on the original open-source benchmark: 82.9% / Score on this Mercor-extended benchmark: 82.7%
Opus 5Max
Score on the original open-source benchmark: 93.5% / Score on this Mercor-extended benchmark: 82.0%
Fable 5Max
Score on the original open-source benchmark: 95.9% / Score on this Mercor-extended benchmark: 78.7%
GPT-5.6 TerraMax
Score on the original open-source benchmark: 77.8% / Score on this Mercor-extended benchmark: 78.7%
Grok 4.6xHigh
Score on the original open-source benchmark: 84.6% / Score on this Mercor-extended benchmark: 78.0%
APEX NEWSLETTER
New benchmarks, leaderboard shifts, and research from the APEX team.
By subscribing you agree to receive updates from Mercor.Unsubscribe anytime.