Get the latest updates about APEX benchmarks – model releases, leaderboard changes, and what they mean for economically valuable work.
APEX is our family of benchmarks that assess a broader shift in how the workforce is evolving by measuring whether AI can complete professional, economically valuable work. Each benchmark tests a different dimension of professional capability. All tasks are built with Mercor experts and leading industry partners.
Long-horizon, cross-application tasks in professional services
GPT-6 AstraxHigh
62.4% ±3.6%
Fable 5.1Max
62.0% ±3.5%
Opus 5Max
60.6% ±3.6%
Fable 5.1High
60.0% ±3.5%
Fable 5Max
59.2% ±3.7%
Work with the team behind APEX — get custom evaluations for your agents and source expert-built training data across 30+ domains.