The AI Productivity Index

The AI Productivity Index (APEX-1) assesses whether frontier models are capable of performing economically valuable tasks across four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD).

The APEX-1 leaderboard

We created APEX-1 to bridge the gap between what professionals want from AI systems and what benchmarks test for. The prompts are realistic, challenging and diverse, and are provided with source documents. Each case was created by a veteran industry expert to capture their day-to-day work.

The leaderboard is based on a hidden heldout set of 400 tasks (n=100 per job). For each task we collect responses from each model 8 times. We grade them using a Judge LM and report the mean value.


To support open research, we have open-sourced n=100 cases that are in-distribution of APEX on Hugging Face, and our eval harness.

Jobs covered in APEX-1

Drafts and reviews contracts, conducts legal research, and advises clients on regulatory and transactional matters. Collaborates with partners on litigation, mergers and acquisitions, and compliance while managing heavy workloads across cases.

Advised by Cass Sunstein—Harvard law professor, former White House Regulatory Administrator, and top-cited legal scholar.

Experts from Latham & Watkins, Skadden, Cravath.

gpt-5-high

GPT-5

76.6%

claude-opus-4-6-high

Opus 4.6

76.4%

claude-opus-4-6-max

Opus 4.6

76.2%

Analyzes industries, evaluates markets, and builds strategic or financial models to guide client decisions. Work often includes preparing presentations, drafting reports, and synthesizing research into actionable recommendations.

Advised by Dominic Barton—former McKinsey Global Managing Director and Canadian Ambassador to China.

Experts from McKinsey, BCG, Deloitte, Accenture, EY.

gpt-6-astra

GPT-6 Astra

78.1%

gpt-5.6-terra-max

GPT-5.6 Terra

76.2%

claude-opus-5-xhigh

Opus 5

73.7%

Diagnoses and treats a wide range of patient conditions, from acute illnesses to chronic diseases. Reviews medical histories, orders and interprets tests, prescribes treatments, and provides preventative care and ongoing patient guidance.

Advised by Eric Topol—Cardiologist, geneticist, and founder of the Scripps Research Translational Institute, leading voice in digital and precision medicine.

Experts from University of Pennsylvania, Northwestern, Cornell, Brigham & Women’s, Mount Sinai.

claude-opus-4-6-max

Opus 4.6

71.6%

claude-opus-4-6-high

Opus 4.6

70.6%

claude-opus-5-xhigh

Opus 5

69.8%

Builds financial models, values companies, and prepares pitch materials for potential deals. Responsibilities include conducting industry research, supporting transaction execution, and producing client-ready presentations under tight deadlines.

Experts from Goldman Sachs, Morgan Stanley, JPMorgan, Barclays, UBS, Bank of America, Evercore.

gpt-5.6-terra-max

GPT-5.6 Terra

71.8%

gpt-6-astra

GPT-6 Astra

70.3%

gemini-3.6-flash

Gemini 3.6 Flash

69.2%

Frequently Asked Questions

The Mercor AI Productivity Index (APEX) is a family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. APEX-1 evaluates single-turn expert knowledge work across high-value sectors, like corporate law, investment banking, management consulting, and medicine.

Rankings can change whenever a frontier model is released. New models are evaluated on APEX-Agents, APEX-SWE, APEX-Accounting and APEX-1 when they ship.

Practicing professionals from leading firms that the tasks simulate, including attorneys from Latham & Watkins, Skadden, and Cravath; consultants from McKinsey and BCG; bankers from Goldman Sachs, Morgan Stanley, and JPMorgan; and physicians from Brigham & Women's, UPenn, and Northwestern.

APEX-1 evaluates the quality of completed work. Model outputs are graded using expert-authored rubrics with an LM judge.

Mean Score measures the average percentage of criteria that are passed for all tasks in the benchmark. It is the primary metric used to rank the APEX-1 leaderboard. Pass@1 measures the percentage of tasks that an agent successfully completes on a single attempt (i.e., scores 100% on the rubric). Together, they capture models’ overall capability and end-to-end task reliability.

Yes, on the open subset. The eval harness is published on GitHub and sample tasks are on HuggingFace, so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.

Yes. Mercor licenses off-the-shelf datasets built by the same expert network. 50,000+ tasks across 30+ domains, with samples available the same day. Learn more about off-the-shelf data.

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.