Real-world code issue resolution.
Tests whether LLMs can resolve real-world GitHub issues by editing codebases, spanning 500 problems from 12 popular Python repos and requiring multi-file, long-context changes.
Model
Score
Fable 5Max
95.9%
Opus 5Max
93.5%
Fable 5.1High
92.1%
Opus 4.8Max
88.5%
Opus 4.7Max
83.1%
APEX NEWSLETTER
New benchmarks, leaderboard shifts, and research from the APEX team.
By subscribing you agree to receive updates from Mercor.Unsubscribe anytime.