Harvey LAB

Developed byHarvey

Built to evaluate and improve agent capabilities for supporting legal work

1,749Public tasks
17.8%Highest score

The Harvey LAB leaderboard

Harvey’s Legal Agent Benchmark (LAB) was created to evaluate agents on the real work of lawyers at law firms and in-house legal teams. The tasks require agents to work through a full client matter, decide which documents matter, and produce finished work products such as memos, redlines, and filings.

LAB was built in three steps. First, practicing lawyers broke real client matters into the tasks a partner would give an associate. Second, they wrote expert rubrics that reflect what a partner or client would check. Third, agents ran each task in a sandbox with the matter files, and a task passes only if every rubric criterion passes.


There are 1,749 tasks in LAB across 24 practice areas and in-house contracting, graded by more than 100,000 rubric criteria. The full LAB dataset is available open-source, along with Harvey's execution harness for running and scoring agents.

Frequently Asked Questions

Harvey's Legal Agent Benchmark (LAB) is an open-source benchmark that tests whether AI agents can complete long, multi-step legal work. Harvey built it. Each task has an instruction, a client matter with the relevant documents, and a required work product, such as a memo, redline, issues list, or filing. Expert rubrics grade that work product.

Harvey built LAB from real client matters handled by practicing lawyers. The team split each matter into the separate tasks a partner would normally give an associate. Practicing attorneys wrote the rubrics to reflect what a supervising partner or client would check before the work goes out.

In one corporate M&A task, the agent gets a virtual data room for a fictional $458M acquisition and a short request from a partner. The data room holds eight material contracts plus other files that may or may not matter, such as a 10-K and a deferred compensation plan. The agent has to find the change-of-control provisions, assess the deal risk, recommend next steps, and draft a memo for the deal team and board. The rubric for that task has 57 criteria covering nine legal issues spread across the documents.

LAB uses all-pass grading. An LLM judge checks the agent's deliverables against each rubric criterion one at a time and returns pass or fail. A task scores 100% only if every criterion passes. Otherwise it scores 0. The leaderboard reports Pass@1: the share of tasks an agent fully completes on a single attempt.

All-pass grading gives no partial credit. In legal work, a diligence memo that misses one material issue can change the economics of a deal or mean the analysis has to be redone. LAB scores reflect how often an agent gets everything right, not how close it came.

Earlier legal benchmarks, such as LegalBench, CUAD, LEXam, and Harvey's own BigLaw Bench, test short-horizon reasoning: read a contract and answer a question, or compare two cases. LAB tests long-horizon work. The agent has to search a full client matter, decide what's relevant, and produce a finished deliverable.

They measure different things. APEX-Agents, developed in collaboration with Harvey and Box, tests agents across three professional services jobs (corporate lawyer, investment banking analyst, and management consultant) using multiple applications. Harvey LAB goes deep on legal work only, with 1,749 tasks across 24 practice areas and in-house contracting.

Rankings can change whenever a new frontier model comes out. Mercor evaluates new models on Harvey LAB and the APEX benchmarks when they launch. See the leaderboard above for current results.

Yes. The full task set, rubrics, and execution harness are on GitHub under the MIT license. The repo includes a step-by-step tutorial that runs one M&A data-room task from setup through scoring.

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.