Oct 8, 2026Company

Why Mercor is building a research team

Mercor
Mercor

Every job in the economy is being reshaped by advances in AI. We’ve seen this with our own APEX benchmarks, where models progressed rapidly in the past year on real-world work.

As they get better, the questions get harder.

How do you measure a model against a job that AI itself is changing? What data actually helps when a model fails at real work? How can human expertise best push the model capability frontier forward?

This is the Mercor Research agenda.

Mercor works with domain experts across thousands of different kinds of work. We use their expertise to build evaluations, understand where models fail, and create data to improve them. Our research team works across that process, from evals through post-training.

Measuring AI on real work

The global economy contains thousands of occupations, each with its own workflows and standards. Very little of that work has a rigorous way to measure what AI can and cannot do.

Our suite of APEX benchmarks cover management consulting, investment banking, corporate law, accounting, and software engineering. We are continuing to expand APEX to understand how AI is changing all types of work throughout the economy.

But the work we are measuring is changing too. A task that matters today may matter less in a year. People may spend less time producing work and more time reviewing it. Different failure modes matter as models move from answering questions to taking actions.

That makes task realism and safety increasingly important. We will measure models in the environments where the work actually happens, including whether they can complete the work reliably without introducing new risks such as exploiting flawed incentives, or concealing mistakes.

Producing better training data

Much of the knowledge models need is spread across where people actually do their work - documents, software tools, email, and the judgment people use while working across it all.

Organizing that knowledge at a useful scale requires more than having experts produce more data. Experts’ time is expensive and limited. One problem we are working on is how to get more useful training signals from each hour they work with us.

An expert might create the task, define the standard for a good deliverable, or act as judge in the handful of cases where models and automated graders disagree.

We also want to make data quality measurable. That means testing agreement between evaluators, validating against held-out examples, and measuring whether a dataset improves model performance.

Closing the loop with post-training

Once we can identify the capabilities it is missing, we can use that signal to train it.

We recently published a post-training guide showing how we increased Qwen3.5-397B’s Pass@1 on APEX-Agents from 16% to 27%, including the training script, model weights, and evaluation traces.

APEX shows us where a model struggles. Experts help us understand why and produce training data around those weaknesses. We post-train the model and measure it again. The information tells us what worked and what remains unsolved.

We want to understand when specialized models outperform general ones, which capabilities respond to targeted post-training, and how the amount and type of expert data affect the result.

Mercor Research team

These problems span AI evaluation, model training, economics, and labor. We are building a team to work on each part.

Our researchers include:

  • Edward J. Hu: first author of LoRA, formerly at OpenAI
  • Minyang Tian: first author of SciCode and CritPt
  • Victor Barres: first author of tau^2-bench and tau-voice, formerly at Sierra
  • Hardik Bhatnagar: co-lead of PostTrainBench
  • Bertie Vidgen: first author of APEX-Agents
  • Along with researchers from Anthropic, Ai2, Bridgewater, DE Shaw

There is still a lot we do not know about how to measure that gap, what data closes it, and how quickly models can learn from the expertise embedded in the economy. Mercor Research is being built to explore these questions at the frontier.