APEX newsletter

Get the latest updates about APEX benchmarks – model releases, leaderboard changes, and what they mean for economically valuable work.

Get the latest AI news – unsubscribe anytime.

What are APEX benchmarks?

APEX is our family of benchmarks that assess a broader shift in how the workforce is evolving by measuring whether AI can complete professional, economically valuable work. Each benchmark tests a different dimension of professional capability. All tasks are built with Mercor experts and leading industry partners.

APEX-Agents

Long-horizon, cross-application tasks in professional services

View
GPT-6 Astra

GPT-6 AstraxHigh

62.4% ±3.6%

Fable 5.1

Fable 5.1Max

62.0% ±3.5%

Opus 5

Opus 5Max

60.6% ±3.6%

Fable 5.1

Fable 5.1High

60.0% ±3.5%

Fable 5

Fable 5Max

59.2% ±3.7%

50%
60%
70%
80%
Mean Score

Want to evaluate models or source training data?

Work with the team behind APEX — get custom evaluations for your agents and source expert-built training data across 30+ domains.

Talk to our team