The AI Productivity Index for Agents (APEX-Agents) measures whether frontier AI agents can execute long-horizon, cross-application tasks across three jobs in professional services.
We created APEX-Agents to evaluate agents on the real day-to-day work of professionals: investment banking analysts, management consultants, and corporate lawyers. The tasks require agents to reason, demonstrate advanced knowledge, use multiple applications, and plan over long horizons.
APEX-Agents was built in three steps. First, industry professionals created a data-rich world over 5-10 days, based on a unique project scenario. Second, they created realistic, challenging tasks using the files from within the world. Third, we gave agents access so they could execute the tasks (with all of the software that a human would use).
There are 33 worlds in APEX-agents, comprising 480 tasks and grading rubrics. The entire APEX-agents dataset is available open-source, along with Archipelago, our infra service for executing and evaluating agent trajectories.
Experts from Latham & Watkins, Skadden, Cravath
View more
GPT-6 Astra
67.9%
Fable 5
67.4%
Opus 5
66.5%
Experts from McKinsey, BCG, Deloitte, Accenture, EY
View more
GPT-6 Astra
67.3%
Fable 5.1
65.4%
Muse Spark 1.1
63.1%
Experts from Goldman Sachs, Morgan Stanley, JPMorgan, Barclays
View more
Fable 5.1
55.7%
Fable 5.1
54.8%
Fable 5
53.2%
The Mercor AI Productivity Index (APEX) is a family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. APEX-Agents evaluates long-horizon, multi-step workflows across high-value sectors, like corporate law, investment banking, and management consulting. It was developed in collaboration with Box and Harvey.
Rankings can change whenever a frontier model is released. New models are evaluated on APEX-Agents, APEX-SWE, APEX-Accounting and APEX-1 when they ship.
Practicing professionals from leading firms that the tasks simulate, including attorneys from Latham & Watkins, Skadden, and Cravath; consultants from McKinsey and BCG; and bankers from Goldman Sachs, Morgan Stanley, and JPMorgan.
APEX-Agents evaluates the quality of completed work. Model outputs are graded using expert-authored rubrics with an LM judge.
Mean Score measures the average percentage of criteria that are passed for all tasks in the benchmark. It is the primary metric used to rank the APEX-Agents leaderboard. Pass@1 measures the percentage of tasks that an agent successfully completes on a single attempt (i.e., scores 100% on the rubric). Together, they capture models’ overall capability and end-to-end task reliability.
Yes, on the open subset. The eval harness is published on GitHub and sample tasks are on HuggingFace, so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.
Yes. Mercor licenses off-the-shelf datasets built by the same expert network. 50,000+ tasks across 30+ domains, with samples available the same day. Learn more about off-the-shelf data.