The AI Productivity Index for Agents

The AI Productivity Index for Agents (APEX-Agents) measures whether frontier AI agents can execute long-horizon, cross-application tasks across three jobs in professional services.

The APEX-Agents leaderboard

We created APEX-Agents to evaluate agents on the real day-to-day work of professionals: investment banking analysts, management consultants, and corporate lawyers. The tasks require agents to reason, demonstrate advanced knowledge, use multiple applications, and plan over long horizons.

APEX-Agents was built in three steps. First, industry professionals created a data-rich world over 5-10 days, based on a unique project scenario. Second, they created realistic, challenging tasks using the files from within the world. Third, we gave agents access so they could execute the tasks (with all of the software that a human would use).


There are 33 worlds in APEX-agents, comprising 480 tasks and grading rubrics. The entire APEX-agents dataset is available open-source, along with Archipelago, our infra service for executing and evaluating agent trajectories.

Jobs evaluated in APEX-Agents

Drafts and reviews contracts, conducts legal research, and advises clients on regulatory and transactional matters. Collaborates with partners on litigation, mergers and acquisitions, and compliance while managing heavy workloads across cases.

Experts from Latham & Watkins, Skadden, Cravath

gpt-6-astra

GPT-6 Astra

67.9%

claude-fable-5

Fable 5

67.4%

claude-opus-5

Opus 5

66.5%

Analyzes industries, evaluates markets, and builds strategic or financial models to guide client decisions. Work often includes preparing presentations, drafting reports, and synthesizing research into actionable recommendations.

Experts from McKinsey, BCG, Deloitte, Accenture, EY

gpt-6-astra

GPT-6 Astra

67.3%

claude-fable-5.1

Fable 5.1

65.4%

muse-spark-1-1

Muse Spark 1.1

63.1%

Builds financial models, values companies, and prepares pitch materials for potential deals. Responsibilities include conducting industry research, supporting transaction execution, and producing client-ready presentations under tight deadlines.

Experts from Goldman Sachs, Morgan Stanley, JPMorgan, Barclays

claude-fable-5.1

Fable 5.1

55.7%

claude-fable-5.1-high

Fable 5.1

54.8%

claude-fable-5

Fable 5

53.2%

Frequently Asked Questions

The Mercor AI Productivity Index (APEX) is a family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. APEX-Agents evaluates long-horizon, multi-step workflows across high-value sectors, like corporate law, investment banking, and management consulting. It was developed in collaboration with Box and Harvey.

Rankings can change whenever a frontier model is released. New models are evaluated on APEX-Agents, APEX-SWE, APEX-Accounting and APEX-1 when they ship.

Practicing professionals from leading firms that the tasks simulate, including attorneys from Latham & Watkins, Skadden, and Cravath; consultants from McKinsey and BCG; and bankers from Goldman Sachs, Morgan Stanley, and JPMorgan.

APEX-Agents evaluates the quality of completed work. Model outputs are graded using expert-authored rubrics with an LM judge.

Mean Score measures the average percentage of criteria that are passed for all tasks in the benchmark. It is the primary metric used to rank the APEX-Agents leaderboard. Pass@1 measures the percentage of tasks that an agent successfully completes on a single attempt (i.e., scores 100% on the rubric). Together, they capture models’ overall capability and end-to-end task reliability.

Yes, on the open subset. The eval harness is published on GitHub and sample tasks are on HuggingFace, so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.

Yes. Mercor licenses off-the-shelf datasets built by the same expert network. 50,000+ tasks across 30+ domains, with samples available the same day. Learn more about off-the-shelf data.

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.