The AI Productivity Index family of benchmarks assesses whether frontier AI models can perform economically valuable tasks across professional services, medicine, and software engineering.
Each benchmark tests a different dimension of professional capability. All tasks are built with Mercor experts and leading industry partners.
Long-horizon, cross-application tasks in professional services
Tests whether AI agents can complete multi-hour professional tasks across investment banking, corporate law, and management consulting, using real tools in Google Suite.
Opus 5Max
60.6% ±3.6%
Fable 5Max
59.2% ±3.7%
Muse Spark 1.1xHigh
58.1% ±3.4%
Long-horizon, cross-application tasks in professional accounting
Measuring AI agents ability to complete professional accounting tasks across tools like accounting software, spreadsheets, and PDFs.
Fable 5Max
56.4% ±3.8%
Muse Spark 1.1xHigh
52.6% ±3.7%
GPT-5.6 SolMax
51.5% ±3.9%
Real-world software engineering across integration and observability
Measures AI performance on real-world software engineering tasks, from bug fixes to feature builds. Built in collaboration with Cognition.
Opus 5Max
63.7% ±6.4%
Fable 5Max
58.8% ±6.4%
Grok 4.6High
56.4% ±6.2%
Single-turn text tasks
Tests whether frontier models can perform economically valuable tasks across professional domains including investment banking, corporate law, management consulting, and medicine.
GPT 5.4High
67.2% ±2.4%
Opus 4.6Max
65.7% ±2.6%
Opus 4.6High
65.3% ±2.7%
Popular open-source benchmarks, independently evaluated by Mercor.
An expert-built benchmark of 80 real lab problems across 16 scientific fields, scored on its 288 test subproblems. Unlike typical coding benchmarks, it pairs domain science with programming skill.
Fable 5Max
49.0%
GPT-5.6 SolxHigh
45.8%
Opus 5Max
45.3%
Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.
GPT-5.5xHigh
83.1%
Kimi K3Max
82.0%
Grok 4.5High
79.0%
Benchmarks coding agents beyond issue resolution across Codebase Q&A (124 tasks), test writing (90), and refactoring (70), combining programmatic validation with rubric-based software-quality scoring (maintainability, abstractions, hygiene). Uses under-specified, agentic task formulations.
Kimi K3Max
52.7%
GLM-5.2Max
46.0%
DeepSeek-V4-FlashMax
44.9%
The latest research and insights from the Mercor team.