Multilingual code issue resolution.
A multilingual expansion for SWE-bench that tests whether LLMs can resolve 298 real-world GitHub issues in 42 repositories spanning nine programming languages.
Model
Score
DeepSeek-V4-Pro-0813Max
96.4%
Opus 5Max
96.2%
DeepSeek-V4-FlashMax
95.4%
Fable 5Max
94.8%
GPT-5.6 TerraMax
94.4%
APEX NEWSLETTER
New benchmarks, leaderboard shifts, and research from the APEX team.
By subscribing you agree to receive updates from Mercor.Unsubscribe anytime.