DeepResearch Bench II

Long-form web research, graded against expert rubrics.

130Public tasks
56.1%Highest score

The DeepResearch Bench II leaderboard

A bilingual benchmark of 130 open-ended research briefs across 22 domains, scored against 9,287 expert-written criteria. Each task is derived from a real review article that the model is explicitly forbidden from consulting, and credit only comes from independently rediscovering the findings.

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.