APEX newsletter

Get the latest updates about APEX benchmarks – model releases, leaderboard changes, and what they mean for economically valuable work.

Get the latest AI news – unsubscribe anytime.

What are APEX benchmarks?

APEX is our family of benchmarks that assess a broader shift in how the workforce is evolving by measuring whether AI can complete professional, economically valuable work. Each benchmark tests a different dimension of professional capability. All tasks are built with Mercor experts and leading industry partners.

APEX-Agents

Long-horizon, cross-application tasks in professional services

View
Opus 5

Opus 5Max

43.5% ±4.2%

Fable 5

Fable 5Max

43.3% ±4.1%

Muse Spark 1.1

Muse Spark 1.1xHigh

41.9% ±3.9%

GPT 5.6 Sol (Max + Pro)

GPT 5.6 Sol (Max + Pro)

40.0% ±4.1%

GPT 5.6 Sol

GPT 5.6 SolMax

39.9% ±4.0%

30%
40%
50%
60%
Pass@1

Want to evaluate models or source training data?

Work with the team behind APEX — get custom evaluations for your agents and source expert-built training data across 30+ domains.

Talk to our team