APEX Benchmarks

The AI Productivity Index family of benchmarks assesses whether frontier AI models can perform economically valuable tasks across professional services, medicine, and software engineering.

Industry partner benchmarks

Benchmarks built with industry leaders, spanning APEX, extended, and open-source evaluations

Legal

In partnership withHarvey

Harvey LAB

1749 Tasks

Attorney-built benchmark of 1,749 real legal tasks across 24 practice areas, graded on finished work products like memos and redlines.

Grok 4.7

Grok 4.7xHigh

19.5% ±4.5%

Muse Spark 1.2

Muse Spark 1.2xHigh

19.4% ±4.6%

Opus 5

Opus 5Max

17.8% ±1.5%

Fable 5.1

Fable 5.1Max

17.4% ±4.2%

Muse Spark 1.3

Muse Spark 1.3Max

15.0% ±4.1%

Accounting

In partnership withRamp

APEX-Accounting

160 Tasks1 Domain

Measuring AI agents ability to complete professional accounting tasks across tools like accounting software, spreadsheets, and PDFs.

Opus 5.5

Opus 5.5Max

61.8% ±4.0%

Fable 5.1

Fable 5.1Max

61.0% ±3.7%

GPT 6 Astra

GPT 6 AstraMax

57.9% ±3.8%

Fable 5

Fable 5Max

56.4% ±3.6%

Opus 5

Opus 5Max

54.5% ±3.9%

Software

In partnership withCognition

FrontierCode

150 Tasks

Maintainer-built benchmark of 150 real open-source coding tasks, graded on whether a patch would merge based on correctness, tests, scope, and conventions.

Opus 5.5

Opus 5.5Medium

54.6% ±0.0%

Fable 5

Fable 5xHigh

53.5% ±0.0%

Opus 5

Opus 5Medium

53.4% ±0.0%

GPT-6 Astra

GPT-6 AstraMax

53.3% ±0.0%

Sonnet 5.5

Sonnet 5.5xHigh

52.1% ±0.0%

Software

In partnership withCognition

APEX-SWE

200 Tasks2 Domains

Measures AI performance on real-world software engineering tasks, from bug fixes to feature builds. Built in collaboration with Cognition.

Opus 5.5

Opus 5.5Max

67.6% ±5.9%

Sonnet 5.5

Sonnet 5.5Max

66.4% ±6.1%

Opus 5

Opus 5Max

63.7% ±6.4%

Fable 5.1

Fable 5.1Max

63.6% ±6.3%

Fable 5

Fable 5Max

58.8% ±6.4%

Banking

In partnership withSierra

τ³-Banking

97 Tasks

Fintech benchmark from Sierra AI that tests whether agents can handle customer-support tasks over 700 interconnected policy documents, graded on real backend outcomes like dispute resolution and account management.

Opus 5.5

Opus 5.5Max

55.0% ±8.2%

Fable 5.1

Fable 5.1Max

54.0% ±8.9%

Sonnet 5.5

Sonnet 5.5Max

54.0% ±8.9%

Muse Spark 1.3

Muse Spark 1.3xHigh

52.2% ±8.6%

Gemini 3.8 Flash

Gemini 3.8 FlashHigh

48.8% ±8.8%

Our benchmarks

Each benchmark tests a different dimension of professional capability. All tasks are built with Mercor experts and leading industry partners.

APEX-Agents

Long-horizon, cross-application tasks in professional services

Tests whether AI agents can complete multi-hour professional tasks across investment banking, corporate law, and management consulting, using real tools in Google Suite.

View Leaderboard
Gemini 4 Argon

Gemini 4 ArgonHigh

82.4% ±4.3%

Sonnet 5.5

Sonnet 5.5Max

75.5% ±4.7%

Opus 5.5

Opus 5.5Max

73.5% ±4.9%

Fable 5.1

Fable 5.1Max

68.6% ±4.9%

Gemini 3.7 Flash

Gemini 3.7 FlashHigh

67.8% ±5.1%

APEX-1

Single-turn text tasks

Tests whether frontier models can perform economically valuable tasks across professional domains including investment banking, corporate law, management consulting, and medicine.

View Leaderboard
Muse Spark 1.2

Muse Spark 1.2xHigh

75.6% ±2.3%

Muse Spark 1.3

Muse Spark 1.3xHigh

74.0% ±2.3%

Fable 5.1

Fable 5.1Max

73.8% ±2.7%

Muse Spark 1.1

Muse Spark 1.1xHigh

73.8% ±2.2%

GPT 5.6 Sol

GPT 5.6 SolMax • Pro

73.1% ±2.4%

Mercor-extended benchmarks

Open-source benchmarks extended with new expert-built Mercor tasks.

100 Mercor tasks130 public tasks

BrowseComp Extended

130 questions requiring agents to persistently navigate the internet to locate hard-to-find, entangled information. Tests information-discovery competence through persistence and creative problem-solving.

Sonnet 5.5

Sonnet 5.5Max

Score on the original open-source benchmark: 86.6% / Score on this Mercor-extended benchmark: 73.0%

GPT 6.1 Sol

GPT 6.1 SolMax

Score on the original open-source benchmark: 95.4% / Score on this Mercor-extended benchmark: 67.7%

Fable 5.1

Fable 5.1Max

Score on the original open-source benchmark: 85.2% / Score on this Mercor-extended benchmark: 63.3%

GPT 6 Sol

GPT 6 SolMax

Score on the original open-source benchmark: 89.4% / Score on this Mercor-extended benchmark: 63.2%

Opus 5.5

Opus 5.5Max

Score on the original open-source benchmark: 88.5% / Score on this Mercor-extended benchmark: 60.5%

More info
150 Mercor tasks5000 public tasks

CharXiv Extended

Evaluates multimodal models on realistic chart understanding using 2,323 hand-curated arXiv charts, with descriptive questions (basic elements) and reasoning questions (synthesis across complex elements).

Muse Spark 1.2

Muse Spark 1.2xHigh

Score on the original open-source benchmark: 95.3% / Score on this Mercor-extended benchmark: 87.1%

Muse Spark 1.1

Muse Spark 1.1xHigh

Score on the original open-source benchmark: 95.2% / Score on this Mercor-extended benchmark: 86.4%

GPT 6 Astra

GPT 6 AstraxHigh

Score on the original open-source benchmark: 95.8% / Score on this Mercor-extended benchmark: 85.1%

GPT 6.1 Sol

GPT 6.1 SolMax

Score on the original open-source benchmark: 96.0% / Score on this Mercor-extended benchmark: 85.1%

Opus 5.5

Opus 5.5Max

Score on the original open-source benchmark: 96.4% / Score on this Mercor-extended benchmark: 84.7%

More info
132 Mercor tasks213 public tasks

GDPval Extended

Evaluates models on real-world, economically valuable tasks spanning most BLS work activities for 44 occupations across nine major GDP sectors.

Opus 5.5

Opus 5.5Max

Score on the original open-source benchmark: 88.3% / Score on this Mercor-extended benchmark: 75.8%

Sonnet 5.5

Sonnet 5.5Max

Score on the original open-source benchmark: 86.5% / Score on this Mercor-extended benchmark: 70.9%

Fable 5.1

Fable 5.1Max

Score on the original open-source benchmark: 86.6% / Score on this Mercor-extended benchmark: 68.7%

Haiku 5.5

Haiku 5.5Max

Score on this Mercor-extended benchmark: 66.0%

Fable 5

Fable 5Max

Score on the original open-source benchmark: 78.6% / Score on this Mercor-extended benchmark: 65.3%

More info
447 Mercor tasks2158 public tasks

HLE Extended

Tests models on 2,158 frontier-level, closed-ended academic questions requiring genuine subject mastery rather than retrieval.

GPT 5.6 Sol

GPT 5.6 SolMax • Pro

Score on this Mercor-extended benchmark: 77.3%

GPT 5.6 Terra

GPT 5.6 TerraMax

Score on the original open-source benchmark: 44.9% / Score on this Mercor-extended benchmark: 70.7%

Opus 5

Opus 5High

Score on the original open-source benchmark: 52.0% / Score on this Mercor-extended benchmark: 61.0%

Fable 5

Fable 5High

Score on the original open-source benchmark: 51.5% / Score on this Mercor-extended benchmark: 58.7%

GPT 5.6 Luna

GPT 5.6 LunaMax

Score on the original open-source benchmark: 39.6% / Score on this Mercor-extended benchmark: 57.5%

More info
89 Mercor tasks100 public tasks

Long Context Reasoning (AA-LCR) Extended

Long Context Reasoning (AA-LCR): 100 hard questions over 234 documents in 30 sets (~100k tokens/question) that require synthesizing information across multiple documents and complex reasoning rather than simple retrieval.

GPT 5.6 Sol

GPT 5.6 SolMax • Pro

Score on the original open-source benchmark: 85.3% / Score on this Mercor-extended benchmark: 79.5%

GLM 5.3

GLM 5.3Max

Score on the original open-source benchmark: 78.0% / Score on this Mercor-extended benchmark: 76.8%

GPT 5.6 Terra

GPT 5.6 TerraMax

Score on the original open-source benchmark: 83.3% / Score on this Mercor-extended benchmark: 74.4%

Opus 5.5

Opus 5.5Max

Score on the original open-source benchmark: 85.0% / Score on this Mercor-extended benchmark: 74.2%

Opus 5

Opus 5Max

Score on the original open-source benchmark: 82.0% / Score on this Mercor-extended benchmark: 73.9%

More info
51 Mercor tasks200 public tasks

MedXpertQA MM Extended

200 multimodal multiple-choice medical questions spread across 11 body systems. Each task requires multi-step reasoning over clinical images and patient details.

Opus 5.5

Opus 5.5Max

Score on the original open-source benchmark: 85.7% / Score on this Mercor-extended benchmark: 71.2%

GPT 6 Sol

GPT 6 SolMax

Score on the original open-source benchmark: 81.3% / Score on this Mercor-extended benchmark: 68.0%

Fable 5.1

Fable 5.1High

Score on the original open-source benchmark: 79.7% / Score on this Mercor-extended benchmark: 66.0%

GPT 5.6 Sol

GPT 5.6 SolMax • Pro

Score on the original open-source benchmark: 79.5% / Score on this Mercor-extended benchmark: 64.7%

GPT 6.1 Sol

GPT 6.1 SolMax

Score on the original open-source benchmark: 85.7% / Score on this Mercor-extended benchmark: 64.7%

More info
140 Mercor tasks1730 public tasks

MMMU-Pro Extended

A more robust multimodal benchmark that filters text-only-solvable questions, expands answer options, and adds vision-only inputs (text embedded in images) to test true joint visual+textual reasoning. Forces models to 'see' and 'read' simultaneously.

GPT 6.1 Sol

GPT 6.1 SolMax

Score on the original open-source benchmark: 85.8% / Score on this Mercor-extended benchmark: 61.3%

GPT 6 Astra

GPT 6 AstraMax

Score on the original open-source benchmark: 87.5% / Score on this Mercor-extended benchmark: 60.0%

GPT 5.6 Sol

GPT 5.6 SolMax • Pro

Score on this Mercor-extended benchmark: 51.9%

GPT 6 Sol

GPT 6 SolMax

Score on the original open-source benchmark: 83.0% / Score on this Mercor-extended benchmark: 51.9%

Opus 5.5

Opus 5.5Max

Score on the original open-source benchmark: 87.4% / Score on this Mercor-extended benchmark: 51.0%

More info
50 Mercor tasks124 public tasks

SWE Atlas - Codebase QnA Extended

Benchmarks coding agents beyond issue resolution across Codebase Q&A (124 tasks), test writing (90), and refactoring (70), combining programmatic validation with rubric-based software-quality scoring (maintainability, abstractions, hygiene). Uses under-specified, agentic task formulations.

Sonnet 5.5

Sonnet 5.5Max

Score on the original open-source benchmark: 50.3% / Score on this Mercor-extended benchmark: 44.7%

Opus 5.5

Opus 5.5Max

Score on the original open-source benchmark: 48.9% / Score on this Mercor-extended benchmark: 43.3%

Opus 5

Opus 5Max

Score on the original open-source benchmark: 50.3% / Score on this Mercor-extended benchmark: 40.0%

Fable 5.1

Fable 5.1Max

Score on the original open-source benchmark: 43.3% / Score on this Mercor-extended benchmark: 38.0%

DeepSeek V4.1 Flash

DeepSeek V4.1 FlashMax

Score on the original open-source benchmark: 53.5% / Score on this Mercor-extended benchmark: 36.7%

More info
50 Mercor tasks500 public tasks

SWE-bench Verified Extended

Tests whether LLMs can resolve real-world GitHub issues by editing codebases, spanning 500 problems from 12 popular Python repos and requiring multi-file, long-context changes.

DeepSeek V4 Pro 0813

DeepSeek V4 Pro 0813Max

Score on this Mercor-extended benchmark: 83.3%

DeepSeek V4.1 Flash

DeepSeek V4.1 FlashMax

Score on the original open-source benchmark: 82.9% / Score on this Mercor-extended benchmark: 82.7%

Opus 5

Opus 5Max

Score on the original open-source benchmark: 93.5% / Score on this Mercor-extended benchmark: 82.0%

Fable 5

Fable 5Max

Score on the original open-source benchmark: 95.9% / Score on this Mercor-extended benchmark: 78.7%

GPT 5.6 Terra

GPT 5.6 TerraMax

Score on the original open-source benchmark: 77.8% / Score on this Mercor-extended benchmark: 78.7%

More info
99 Mercor tasks89 public tasks

Terminal-Bench 2.1 Extended

Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.

Opus 5.5

Opus 5.5Max

Score on the original open-source benchmark: 83.5% / Score on this Mercor-extended benchmark: 47.5%

GPT 6.1 Sol

GPT 6.1 SolMax

Score on the original open-source benchmark: 85.4% / Score on this Mercor-extended benchmark: 45.8%

Grok 4.6

Grok 4.6xHigh

Score on the original open-source benchmark: 84.6% / Score on this Mercor-extended benchmark: 45.5%

Sonnet 5.5

Sonnet 5.5Max

Score on the original open-source benchmark: 87.3% / Score on this Mercor-extended benchmark: 45.5%

Gemini 3.8 Flash

Gemini 3.8 FlashHigh

Score on the original open-source benchmark: 80.9% / Score on this Mercor-extended benchmark: 44.4%

More info
New
Off-the-shelf data
License-ready datasets, expert-written and graded. 50K+ tasks across 30+ domains, ready to train on today.

Open-source benchmarks

Popular open-source benchmarks, independently evaluated by Mercor.

65 public tasks

SciCode

An expert-built benchmark of 80 real lab problems across 16 scientific fields, scored on its 288 test subproblems. Unlike typical coding benchmarks, it pairs domain science with programming skill.

Opus 5.5

Opus 5.5Max

57.4%

Fable 5.1

Fable 5.1Max

54.6%

Sonnet 5.5

Sonnet 5.5Max

51.2%

GPT 6 Astra

GPT 6 AstraMax

49.3%

Fable 5

Fable 5Max

49.0%

More info
89 public tasks

Terminal-Bench 2.1

Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.

Sonnet 5.5

Sonnet 5.5Max

87.3%

DeepSeek V4.1 Flash

DeepSeek V4.1 FlashMax

86.5%

GPT 6 Astra

GPT 6 AstraMax

86.1%

GPT 6 Sol

GPT 6 SolMax

85.4%

GPT 6.1 Sol

GPT 6.1 SolMax

85.4%

More info
124 public tasks

SWE Atlas - Codebase QnA

Benchmarks coding agents beyond issue resolution across Codebase Q&A (124 tasks), test writing (90), and refactoring (70), combining programmatic validation with rubric-based software-quality scoring (maintainability, abstractions, hygiene). Uses under-specified, agentic task formulations.

DeepSeek V4.1 Flash

DeepSeek V4.1 FlashMax

53.5%

Kimi K3

Kimi K3Max

52.7%

Qwen 3.8 Max

Qwen 3.8 MaxxHigh

50.5%

Opus 5

Opus 5Max

50.3%

Sonnet 5.5

Sonnet 5.5Max

50.3%

More info
500 public tasks

SWE-bench Verified

Tests whether LLMs can resolve real-world GitHub issues by editing codebases, spanning 500 problems from 12 popular Python repos and requiring multi-file, long-context changes.

Opus 5.5

Opus 5.5Max

98.2%

Sonnet 5.5

Sonnet 5.5Max

96.7%

Fable 5.1

Fable 5.1Max

96.6%

Fable 5

Fable 5Max

95.9%

Opus 5

Opus 5Max

93.5%

More info
100 public tasks

Long Context Reasoning (AA-LCR v1.1)

Long Context Reasoning (AA-LCR): 100 hard questions over 234 documents in 30 sets (~100k tokens/question) that require synthesizing information across multiple documents and complex reasoning rather than simple retrieval.

Fable 5.1

Fable 5.1Max

87.5%

Fable 5

Fable 5Max

85.5%

GPT 5.6 Sol

GPT 5.6 SolMax • Pro

85.3%

Opus 5.5

Opus 5.5Max

85.0%

Kimi K3

Kimi K3Max

84.3%

More info
12032 public tasks

MMLU-Pro

An enhanced MMLU adding harder, reasoning-focused questions and expanding choices from 4 to 10 options. The MMLU benchmark evaluates an AI model's general knowledge and reasoning skills using multiple-choice questions across 57 academic and professional subjects.

Opus 5.5

Opus 5.5Max

93.4%

Fable 5.1

Fable 5.1Max

93.0%

Opus 5

Opus 5Max

92.0%

GPT 6 Astra

GPT 6 AstraMax

91.5%

Gemini 3.1 Pro

Gemini 3.1 ProHigh

91.3%

More info
1730 public tasks

MMMU-Pro

A more robust multimodal benchmark that filters text-only-solvable questions, expands answer options, and adds vision-only inputs (text embedded in images) to test true joint visual+textual reasoning. Forces models to 'see' and 'read' simultaneously.

GPT 6 Astra

GPT 6 AstraMax

87.5%

Opus 5.5

Opus 5.5Max

87.4%

Fable 5.1

Fable 5.1Max

86.1%

Sonnet 5.5

Sonnet 5.5Max

85.8%

GPT 6.1 Sol

GPT 6.1 SolMax

85.8%

More info
198 public tasks

GPQA Diamond

Graduate-level, 'Google-proof' multiple-choice science questions written by domain experts.

GPT 6 Astra

GPT 6 AstraxHigh

96.2%

Sonnet 5.5

Sonnet 5.5Max

95.2%

GPT 6.1 Sol

GPT 6.1 SolMax

95.1%

GPT 5.6 Sol

GPT 5.6 SolMax • Pro

94.9%

GPT 5.5

GPT 5.5xHigh

94.6%

More info
1645 public tasks

AdvancedIF

Evaluates complex, multi-turn, and system-level instruction following via expert-curated rubrics over 1,600+ prompts. Paired with a reinforcement-learning method for improving instruction following.

Gemini 3.1 Pro

Gemini 3.1 ProHigh

86.7%

Gemini 3.6 Flash

Gemini 3.6 FlashHigh

85.3%

Qwen 3.8 Max

Qwen 3.8 MaxxHigh

85.1%

Fable 5.1

Fable 5.1Max

84.2%

DeepSeek V4.1 Flash

DeepSeek V4.1 FlashMax

84.0%

More info
30 public tasks

AIME 2025

Competition-mathematics benchmark drawn from the 2025 American Invitational Mathematics Examination; each answer is an integer 0-999. Measures advanced multi-step mathematical problem-solving and reasoning.

Opus 5

Opus 5Max

100.0%

Fable 5

Fable 5Max

100.0%

DeepSeek V4 Flash

DeepSeek V4 FlashMax

100.0%

GPT 5.5

GPT 5.5xHigh

100.0%

GPT 5.6 Sol

GPT 5.6 SolMax • Pro

100.0%

More info
130 public tasks

BrowseComp

130 questions requiring agents to persistently navigate the internet to locate hard-to-find, entangled information. Tests information-discovery competence through persistence and creative problem-solving.

GPT 5.6 Sol

GPT 5.6 SolMax • Pro

95.7%

GPT 6.1 Sol

GPT 6.1 SolMax

95.4%

GPT 6 Astra

GPT 6 AstraxHigh

94.2%

GPT 6 Sol

GPT 6 SolMax

89.4%

Kimi K3

Kimi K3Max

89.0%

More info
5000 public tasks

CharXiv

Evaluates multimodal models on realistic chart understanding using 2,323 hand-curated arXiv charts, with descriptive questions (basic elements) and reasoning questions (synthesis across complex elements).

Opus 5.5

Opus 5.5Max

96.4%

GPT 6 Astra

GPT 6 AstraMax

96.1%

Sonnet 5.5

Sonnet 5.5Max

96.1%

GPT 6.1 Sol

GPT 6.1 SolMax

96.0%

Gemini 3.8 Flash

Gemini 3.8 FlashHigh

95.4%

More info
517 public tasks

CVDP

517 RTL design and verification problems from real hardware workflows — spec-to-RTL, code completion and modification, module reuse, debugging, testbench generation, and RTL/testbench comprehension Q&A.

Opus 5.5

Opus 5.5Max

64.1%

Fable 5.1

Fable 5.1Max

62.6%

Opus 5

Opus 5Max

62.2%

Sonnet 5.5

Sonnet 5.5Max

60.6%

Grok 4.6

Grok 4.6xHigh

58.0%

More info
130 public tasks

DeepResearch Bench II

A bilingual benchmark of 130 open-ended research briefs across 22 domains, scored against 9,287 expert-written criteria. Each task is derived from a real review article that the model is explicitly forbidden from consulting, and credit only comes from independently rediscovering the findings.

Opus 5

Opus 5High

56.1%

Fable 5.1

Fable 5.1Max

55.0%

Sonnet 4.6

Sonnet 4.6High

53.3%

Opus 5.5

Opus 5.5Max

53.3%

DeepSeek V4.1 Flash

DeepSeek V4.1 FlashMax

53.2%

More info
113 public tasks

DeepSWE v1.1

113 original software engineering tasks spanning 91 repositories and 5 languages. Each task requires large, novel fixes not sourced from existing public commits.

Opus 5.5

Opus 5.5Max

72.3%

GPT 6.1 Sol

GPT 6.1 SolMax

72.3%

GPT 6 Astra

GPT 6 AstraMax

72.0%

DeepSeek V4.1 Flash

DeepSeek V4.1 FlashMax

71.7%

Opus 5

Opus 5Max

71.4%

More info
34 public tasks

FrontierSWE v2

34 open-ended engineering projects, from building compilers and simulators to optimizing inference and training research models. Each task gives an agent 20 hours and scores its result on the task's own metric, with partial credit.

Fable 5.1

Fable 5.1High

37.2%

GPT 6 Astra

GPT 6 AstraMax

29.9%

Opus 5

Opus 5High

27.3%

GPT 5.6 Sol

GPT 5.6 SolMax

20.2%

GLM 5.3

GLM 5.3Max

17.1%

More info
213 public tasks

GDPval

Evaluates models on real-world, economically valuable tasks spanning most BLS work activities for 44 occupations across nine major GDP sectors.

Opus 5.5

Opus 5.5Max

88.3%

GPT 6 Astra

GPT 6 AstraMax

86.7%

DeepSeek V4.1 Flash

DeepSeek V4.1 FlashMax

86.7%

Fable 5.1

Fable 5.1Max

86.6%

Sonnet 5.5

Sonnet 5.5Max

86.5%

More info
10 public tasks

GeneBench-Pro

GeneBench-Pro tests whether AI agents can take messy genomics datasets through QC, modeling and diagnostics to a verifiable quantitative answer. Scores use the 10-problem public set, with 10 runs per model on each problem.

Sonnet 5.5

Sonnet 5.5Max

39.0%

GPT 6 Astra

GPT 6 AstraMax

36.0%

GPT 6.1 Sol

GPT 6.1 SolMax

33.0%

Opus 5.5

Opus 5.5Max

30.0%

Muse Spark 1.3

Muse Spark 1.3Max

26.0%

More info
2158 public tasks

HLE

Tests models on 2,158 frontier-level, closed-ended academic questions requiring genuine subject mastery rather than retrieval.

Muse Spark 1.1

Muse Spark 1.1xHigh

57.9%

Opus 5.5

Opus 5.5Max

57.4%

GPT 6 Astra

GPT 6 AstraxHigh

55.6%

Fable 5.1

Fable 5.1Max

54.6%

GPT 6.1 Sol

GPT 6.1 SolMax

54.5%

More info
200 public tasks

MedXpertQA MM

200 multimodal multiple-choice medical questions spread across 11 body systems. Each task requires multi-step reasoning over clinical images and patient details.

GPT 6 Astra

GPT 6 AstraMax

86.8%

Opus 5.5

Opus 5.5Max

85.7%

GPT 6.1 Sol

GPT 6.1 SolMax

85.7%

Gemini 3.8 Flash

Gemini 3.8 FlashHigh

84.7%

Fable 5.1

Fable 5.1Max

83.7%

More info
100 public tasks

ProgramBench

Tests whether software-engineering agents can rebuild complete programs from scratch given only a program and its documentation, matching a reference executable via end-to-end testing (100 tasks, CLI tools to FFmpeg/SQLite/PHP).

Opus 5.5

Opus 5.5Max

20.0%

Fable 5.1

Fable 5.1High

11.3%

GPT 6 Astra

GPT 6 AstraMax

8.0%

GPT 6.1 Sol

GPT 6.1 SolMax

6.3%

Opus 5

Opus 5High

4.3%

More info
140 public tasks

RedlineBench

Collection of 140 tasks spanning 3 multi-turn negotiation scenarios (2 SaaS MSAs and 1 professional-services MSA) over 4 alternating turns.

GPT 5.6 Sol

GPT 5.6 SolMax

46.8%

GPT 5.5

GPT 5.5xHigh

44.1%

GPT 5.6 Terra

GPT 5.6 TerraMax

44.0%

Opus 5

Opus 5Max

44.0%

GPT 5.6 Luna

GPT 5.6 LunaMax

43.3%

More info
500 public tasks

SUPERChem

500 expert-prepared advanced chemistry problems. Each is scored on its reasoning path against an expert-written solution, not just on the final answer.

GPT 6 Astra

GPT 6 AstraxHigh

80.3%

Opus 5.5

Opus 5.5Max

79.6%

GPT 6.1 Sol

GPT 6.1 SolMax

77.8%

Sonnet 5.5

Sonnet 5.5Max

75.4%

GPT 5.6 Sol

GPT 5.6 SolMax • Pro

74.9%

More info
298 public tasks

SWE-bench Multilingual

A multilingual expansion for SWE-bench that tests whether LLMs can resolve 298 real-world GitHub issues in 42 repositories spanning nine programming languages.

DeepSeek V4.1 Flash

DeepSeek V4.1 FlashMax

98.2%

DeepSeek V4 Pro 0813

DeepSeek V4 Pro 0813Max

96.4%

Opus 5

Opus 5Max

96.2%

Opus 5.5

Opus 5.5Max

95.5%

DeepSeek V4 Flash

DeepSeek V4 FlashMax

95.4%

More info
66 public tasks

Terminal-Bench 4.0

66 terminal tasks spanning frontier engineering and science that require end-to-end autonomous execution and tool use over a long horizon.

GPT 6.1 Sol

GPT 6.1 SolMax

59.6%

Opus 5.5

Opus 5.5Max

58.3%

GPT 6 Astra

GPT 6 AstraMax

54.5%

Fable 5.1

Fable 5.1Max

53.0%

Sonnet 5.5

Sonnet 5.5Max

51.5%

More info

Benchmarking methodology

How Mercor runs benchmarks to measure the frontier of intelligence.

Frequently Asked Questions

The Mercor AI Productivity Index (APEX) is a family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. It provides data-driven, real-world productivity metrics across high-value sectors, like software engineering, corporate law, investment banking, accounting, and management consulting. The suite includes benchmarks such as APEX-Agents, which evaluates long-horizon, multi-step agent workflows; APEX-SWE, focused on software engineering; APEX-Accounting, focused on agentic accounting tasks; and APEX-1, which evaluates single-turn expert knowledge work. Additional benchmarks will be introduced as APEX expands into new domains, data types, and workflows.

Rankings can change whenever a frontier model is released. New models are evaluated on APEX-Agents, APEX-SWE, APEX-Accounting and APEX-1 when they ship.

Practicing professionals from leading firms that the tasks simulate, including attorneys from Latham & Watkins, Skadden, and Cravath; consultants from McKinsey and BCG; bankers from Goldman Sachs, Morgan Stanley, and JPMorgan; and physicians from Brigham & Women's, UPenn, and Northwestern.

Yes, on the open subset. The eval harness is published on GitHub and sample tasks are on HuggingFace, so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.

Frontier labs and model developers can request evaluation using this form.

Yes. Mercor licenses off-the-shelf datasets built by the same expert network. 50,000+ tasks across 30+ domains, with samples available the same day. Learn more about off-the-shelf data.

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.