Long-form web research, graded against expert rubrics.
A bilingual benchmark of 130 open-ended research briefs across 22 domains, scored against 9,287 expert-written criteria. Each task is derived from a real review article that the model is explicitly forbidden from consulting, and credit only comes from independently rediscovering the findings.