The AI Productivity Index family of benchmarks assesses whether frontier AI models can perform economically valuable tasks across professional services, medicine, and software engineering.
Benchmarks built with industry leaders, spanning APEX, extended, and open-source evaluations
Legal
Open-source benchmark that evaluates agent capabilities for supporting legal work.
Attorney-built benchmark of 1,749 real legal tasks across 24 practice areas, graded on finished work products like memos and redlines.
Grok 4.7xHigh
19.5% ±4.5%
Muse Spark 1.2xHigh
19.4% ±4.6%
Opus 5Max
17.8% ±1.5%
Fable 5.1Max
17.4% ±4.2%
Muse Spark 1.3Max
15.0% ±4.1%
Accounting
Long-horizon, cross-application tasks in professional accounting
Measuring AI agents ability to complete professional accounting tasks across tools like accounting software, spreadsheets, and PDFs.
Opus 5.5Max
61.8% ±4.0%
Fable 5.1Max
61.0% ±3.7%
GPT 6 AstraMax
57.9% ±3.8%
Fable 5Max
56.4% ±3.6%
Opus 5Max
54.5% ±3.9%
Software
Open-source benchmark testing models against the standards of production codebases.
Maintainer-built benchmark of 150 real open-source coding tasks, graded on whether a patch would merge based on correctness, tests, scope, and conventions.
Opus 5.5Medium
54.6% ±0.0%
Fable 5xHigh
53.5% ±0.0%
Opus 5Medium
53.4% ±0.0%
GPT-6 AstraMax
53.3% ±0.0%
Sonnet 5.5xHigh
52.1% ±0.0%
Software
Real-world software engineering across integration and observability
Measures AI performance on real-world software engineering tasks, from bug fixes to feature builds. Built in collaboration with Cognition.
Opus 5.5Max
67.6% ±5.9%
Sonnet 5.5Max
66.4% ±6.1%
Opus 5Max
63.7% ±6.4%
Fable 5.1Max
63.6% ±6.3%
Fable 5Max
58.8% ±6.4%
Banking
Open-source fintech customer-support benchmark testing knowledge-intensive agents.
Fintech benchmark from Sierra AI that tests whether agents can handle customer-support tasks over 700 interconnected policy documents, graded on real backend outcomes like dispute resolution and account management.
Opus 5.5Max
55.0% ±8.2%
Fable 5.1Max
54.0% ±8.9%
Sonnet 5.5Max
54.0% ±8.9%
Muse Spark 1.3xHigh
52.2% ±8.6%
Gemini 3.8 FlashHigh
48.8% ±8.8%
Each benchmark tests a different dimension of professional capability. All tasks are built with Mercor experts and leading industry partners.
Long-horizon, cross-application tasks in professional services
Tests whether AI agents can complete multi-hour professional tasks across investment banking, corporate law, and management consulting, using real tools in Google Suite.
Gemini 4 ArgonHigh
82.4% ±4.3%
Sonnet 5.5Max
75.5% ±4.7%
Opus 5.5Max
73.5% ±4.9%
Fable 5.1Max
68.6% ±4.9%
Gemini 3.7 FlashHigh
67.8% ±5.1%
Single-turn text tasks
Tests whether frontier models can perform economically valuable tasks across professional domains including investment banking, corporate law, management consulting, and medicine.
Muse Spark 1.2xHigh
75.6% ±2.3%
Muse Spark 1.3xHigh
74.0% ±2.3%
Fable 5.1Max
73.8% ±2.7%
Muse Spark 1.1xHigh
73.8% ±2.2%
GPT 5.6 SolMax • Pro
73.1% ±2.4%
Open-source benchmarks extended with new expert-built Mercor tasks.
Popular open-source benchmarks, independently evaluated by Mercor.
The latest research and insights from the Mercor team.
How Mercor runs benchmarks to measure the frontier of intelligence.
The Mercor AI Productivity Index (APEX) is a family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. It provides data-driven, real-world productivity metrics across high-value sectors, like software engineering, corporate law, investment banking, accounting, and management consulting. The suite includes benchmarks such as APEX-Agents, which evaluates long-horizon, multi-step agent workflows; APEX-SWE, focused on software engineering; APEX-Accounting, focused on agentic accounting tasks; and APEX-1, which evaluates single-turn expert knowledge work. Additional benchmarks will be introduced as APEX expands into new domains, data types, and workflows.
Rankings can change whenever a frontier model is released. New models are evaluated on APEX-Agents, APEX-SWE, APEX-Accounting and APEX-1 when they ship.
Practicing professionals from leading firms that the tasks simulate, including attorneys from Latham & Watkins, Skadden, and Cravath; consultants from McKinsey and BCG; bankers from Goldman Sachs, Morgan Stanley, and JPMorgan; and physicians from Brigham & Women's, UPenn, and Northwestern.
Yes, on the open subset. The eval harness is published on GitHub and sample tasks are on HuggingFace, so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.
Frontier labs and model developers can request evaluation using this form.
Yes. Mercor licenses off-the-shelf datasets built by the same expert network. 50,000+ tasks across 30+ domains, with samples available the same day. Learn more about off-the-shelf data.