Diverse datasets spanning professional, coding, research, agentic, and multimodal tasks.
No waiting, the same datasets used by the world's top AI labs, ready to license immediately.
Every dataset is pre-built, peer-reviewed, and ready to license. Sample tasks delivered same day. No custom pipeline, no 6-month wait.
Every task is written by PhD researchers, practicing lawyers, and senior engineers, not crowdsourced.
Each task is created, vetted, and peer reviewed by domain experts and using quality automation to ensure training signal.
Every dataset is built from tasks created and evaluated by our expert network.
Mercor's AI Productivity Index for Agents. Expert-built tasks run inside high-fidelity enterprise app clones, testing whether agents can navigate hundreds of files, hold context, and finish long-horizon professional work.
Domains
Featured Datasets
A harder, contamination-resistant successor to SWE-bench across seven languages. Docker-containerized, CI-verified tasks test whether agents can resolve real GitHub issues with multi-file patches.
Mercor's expansion of Terminal-Bench, built to stay hard as models improve. Repository-scale tasks demand multi-phase terminal workflows, with integrity checks that catch destructive or partial work.
A web-browsing benchmark for realistic, end-to-end search. Tasks test whether models can strategize, navigate authoritative sources, and synthesize grounded answers to questions general knowledge can't solve.
Mercor’s off-the-shelf datasets are ready-to-license datasets for AI post-training. The catalog spans professional knowledge work, coding, research, agentic, search, multimodal tasks, and more, with datasets available immediately and many extendable through custom data projects.
Mercor offers datasets across enterprise and professional work, coding, search and browsing, research, STEM, multimodal reasoning, medicine, and consumer agent tasks. Offerings include datasets such as APEX-Agents, APEX-Accounting, GDPVal, BrowseComp, HLE, MMMU Expert, SWE Bench Extension, Terminal Bench Pro, APEX-SWE, and other specialized datasets.
Mercor datasets are built with domain experts and experienced practitioners whose backgrounds match the work being represented. Dataset tasks may include expert-written prompts, rubrics, golden solutions, realistic workplace environments, production engineering environments, long-horizon agent workflows, and more.
Mercor’s datasets are designed around realistic work rather than generic or templated examples. The catalog includes real-world professional tasks, long-horizon agent workflows, repository-scale coding work, open-web research, and multimodal reasoning, giving model developers data that more closely reflects the workloads models are expected to handle in production.
Mercor datasets can be used for model training under a variety of different methods, including SFT, DPO, GRPO, self-d. Different datasets are designed to teach different capabilities, including professional reasoning, instruction following, agentic tool use, software engineering, research, browsing, long-context reasoning, and multimodal understanding.
Yes. Mercor provides samples for its off-the-shelf datasets so AI labs and model developers can evaluate the data before purchasing. Many datasets are immediately available for licensing.
Yes. Mercor licenses off-the-shelf datasets to AI labs and enterprise model developers for model training and related development use cases. Available datasets vary by domain, task type, and volume.
Yes. Many of Mercor’s off-the-shelf datasets can be extended through custom data projects. Mercor can increase volume or build additional data around a customer’s specific domain, capability, or training objective.
Yes. Mercor can source and build custom datasets tailored to specific model-development needs, and its data acquisition team can begin sourcing new datasets within days of a request. Custom projects can support areas such as agent environments, domain-specific reasoning, legal workflows, engineering workflows, and new verticals.
Yes, Mercor offers a full-service research partnership to strategic customers. Services include ongoing model evaluations, failure mode analyses, and various post-training capabilities. Reach out to our team to learn more.
Sample tasks from any dataset, delivered same day.