APEX Benchmarks

The AI Productivity Index family of benchmarks assesses whether frontier AI models can perform economically valuable tasks across professional services, medicine, and software engineering.

Our benchmarks

Each benchmark tests a different dimension of professional capability. All tasks are built with Mercor experts and leading industry partners.

APEX-Agents

Long-horizon, cross-application tasks in professional services

Tests whether AI agents can complete multi-hour professional tasks across investment banking, corporate law, and management consulting, using real tools in Google Suite.

View Leaderboard
Opus 5

Opus 5Max

60.6% ±3.6%

Fable 5

Fable 5Max

59.2% ±3.7%

Muse Spark 1.1

Muse Spark 1.1xHigh

58.1% ±3.4%

APEX-Accounting

Long-horizon, cross-application tasks in professional accounting

Measuring AI agents ability to complete professional accounting tasks across tools like accounting software, spreadsheets, and PDFs.

View Leaderboard
Fable 5

Fable 5Max

56.4% ±3.8%

Muse Spark 1.1

Muse Spark 1.1xHigh

52.6% ±3.7%

GPT-5.6 Sol

GPT-5.6 SolMax

51.5% ±3.9%

APEX-SWE

Real-world software engineering across integration and observability

Measures AI performance on real-world software engineering tasks, from bug fixes to feature builds. Built in collaboration with Cognition.

View Leaderboard
Opus 5

Opus 5Max

63.7% ±6.4%

Fable 5

Fable 5Max

58.8% ±6.4%

Grok 4.6

Grok 4.6High

56.4% ±6.2%

APEX

Single-turn text tasks

Tests whether frontier models can perform economically valuable tasks across professional domains including investment banking, corporate law, management consulting, and medicine.

View Leaderboard
GPT 5.4

GPT 5.4High

67.2% ±2.4%

Opus 4.6

Opus 4.6Max

65.7% ±2.6%

Opus 4.6

Opus 4.6High

65.3% ±2.7%

New
Off-the-shelf data
License-ready datasets, expert-written and graded. 50K+ tasks across 30+ domains, ready to train on today.

Open-source benchmarks

Popular open-source benchmarks, independently evaluated by Mercor.

65 public tasks

SciCode

An expert-built benchmark of 80 real lab problems across 16 scientific fields, scored on its 288 test subproblems. Unlike typical coding benchmarks, it pairs domain science with programming skill.

Fable 5

Fable 5Max

49.0%

GPT-5.6 Sol

GPT-5.6 SolxHigh

45.8%

Opus 5

Opus 5Max

45.3%

More info
89 public tasks

Terminal-Bench 2.1

Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.

GPT-5.5

GPT-5.5xHigh

83.1%

Kimi K3

Kimi K3Max

82.0%

Grok 4.5

Grok 4.5High

79.0%

More info
124 public tasks

SWE Atlas - Codebase QnA

Benchmarks coding agents beyond issue resolution across Codebase Q&A (124 tasks), test writing (90), and refactoring (70), combining programmatic validation with rubric-based software-quality scoring (maintainability, abstractions, hygiene). Uses under-specified, agentic task formulations.

Kimi K3

Kimi K3Max

52.7%

GLM-5.2

GLM-5.2Max

46.0%

DeepSeek-V4-Flash

DeepSeek-V4-FlashMax

44.9%

More info
500 public tasks

SWE-bench Verified

Tests whether LLMs can resolve real-world GitHub issues by editing codebases, spanning 2,294 problems from 12 popular Python repos and requiring multi-file, long-context changes.

Fable 5

Fable 5Max

93.5%

Opus 5

Opus 5Max

93.5%

Grok 4.5

Grok 4.5High

81.7%

More info

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.