The AI Productivity Index for Software Engineers (APEX-SWE) measures whether frontier AI systems can execute economically valuable software engineering work. It covers Integration and Observability tasks.
We created APEX-SWE to evaluate the real day-to-day work of software engineers, unlike unit-level and single-repository bug-fix benchmarks. It comprises n=200 cases and spans two complementary settings that mirror professional SWE work: (1) Integration tasks that require end-to-end system construction and deployment across heterogeneous services and (2) Observability tasks that require debugging with production-style telemetry.
Each task includes a human-authored rubric that grades agent outputs for functional requirements, robustness, and code style, alongside unit tests.
To support open research, we have open-sourced n=50 cases that are in-distribution of APEX-SWE on Hugging Face with all metadata labels. We have also shared our eval harness for reproducibility.
View more
Fable 5.1
68.1%
Gemini 3.6 Flash
64.2%
Opus 5
64.0%
View more
Opus 5
63.5%
Fable 5.1
59.0%
Fable 5
54.2%
The Mercor AI Productivity Index (APEX) is a family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. APEX-SWE evaluates the real day-to-day work of software engineers, including integration tasks that require end-to-end system construction and deployment across heterogeneous services, and observability tasks that require debugging with production-style telemetry. It was developed in collaboration with Cognition.
Rankings can change whenever a frontier model is released. New models are evaluated on APEX-Agents, APEX-SWE, APEX-Accounting and APEX-1 when they ship.
Practicing professional software engineers from Mercor’s network of 5M+ experts.
APEX-SWE evaluates the quality of completed work. Model outputs are graded using expert-authored rubrics with an LM judge.
Pass@1 measures the percentage of tasks that an agent successfully completes on a single attempt (i.e., scores 100% on the rubric). It is the primary metric used to rank the APEX-Agents leaderboard.
Yes, on the open subset. The eval harness is published on GitHub and sample tasks are on HuggingFace, so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.
Yes. Mercor licenses off-the-shelf datasets built by the same expert network. 50,000+ tasks across 30+ domains, with samples available the same day. Learn more about off-the-shelf data.