The AI Productivity Index (APEX-1) assesses whether frontier models are capable of performing economically valuable tasks across four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD).
We created APEX-1 to bridge the gap between what professionals want from AI systems and what benchmarks test for. The prompts are realistic, challenging and diverse, and are provided with source documents. Each case was created by a veteran industry expert to capture their day-to-day work.
The leaderboard is based on a hidden heldout set of 400 tasks (n=100 per job). For each task we collect responses from each model 8 times. We grade them using a Judge LM and report the mean value.
To support open research, we have open-sourced n=100 cases that are in-distribution of APEX on Hugging Face, and our eval harness.
Advised by Cass Sunstein—Harvard law professor, former White House Regulatory Administrator, and top-cited legal scholar.
Experts from Latham & Watkins, Skadden, Cravath.
View more
GPT-5
76.6%
Opus 4.6
76.4%
Opus 4.6
76.2%
Advised by Dominic Barton—former McKinsey Global Managing Director and Canadian Ambassador to China.
Experts from McKinsey, BCG, Deloitte, Accenture, EY.
View more
GPT-6 Astra
78.1%
GPT-5.6 Terra
76.2%
Opus 5
73.7%
Advised by Eric Topol—Cardiologist, geneticist, and founder of the Scripps Research Translational Institute, leading voice in digital and precision medicine.
Experts from University of Pennsylvania, Northwestern, Cornell, Brigham & Women’s, Mount Sinai.
View more
Opus 4.6
71.6%
Opus 4.6
70.6%
Opus 5
69.8%
Experts from Goldman Sachs, Morgan Stanley, JPMorgan, Barclays, UBS, Bank of America, Evercore.
View more
GPT-5.6 Terra
71.8%
GPT-6 Astra
70.3%
Gemini 3.6 Flash
69.2%
The Mercor AI Productivity Index (APEX) is a family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. APEX-1 evaluates single-turn expert knowledge work across high-value sectors, like corporate law, investment banking, management consulting, and medicine.
Rankings can change whenever a frontier model is released. New models are evaluated on APEX-Agents, APEX-SWE, APEX-Accounting and APEX-1 when they ship.
Practicing professionals from leading firms that the tasks simulate, including attorneys from Latham & Watkins, Skadden, and Cravath; consultants from McKinsey and BCG; bankers from Goldman Sachs, Morgan Stanley, and JPMorgan; and physicians from Brigham & Women's, UPenn, and Northwestern.
APEX-1 evaluates the quality of completed work. Model outputs are graded using expert-authored rubrics with an LM judge.
Mean Score measures the average percentage of criteria that are passed for all tasks in the benchmark. It is the primary metric used to rank the APEX-1 leaderboard. Pass@1 measures the percentage of tasks that an agent successfully completes on a single attempt (i.e., scores 100% on the rubric). Together, they capture models’ overall capability and end-to-end task reliability.
Yes, on the open subset. The eval harness is published on GitHub and sample tasks are on HuggingFace, so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.
Yes. Mercor licenses off-the-shelf datasets built by the same expert network. 50,000+ tasks across 30+ domains, with samples available the same day. Learn more about off-the-shelf data.