A large language model (LLM) can produce a polished response but still fail at the work your team needs to complete. A contract summary, for example, may sound convincing while overlooking a clause the reviewer needs to see.
That’s where evaluation comes in. As organizations rely on LLMs for more complex and consequential work, they need a way to determine whether strong-looking outputs reflect reliable performance.
What is LLM evaluation?
LLM evaluation is the process of testing how well a large language model performs on specific tasks. It uses representative data, defined criteria, and measurable outcomes to assess whether a model or AI application is accurate, reliable, and fit for its intended use under realistic working conditions.
Rather than judging a model based on a few impressive responses, LLM evals test performance systematically. These AI evaluations may use automated benchmarks, human reviewers, or both to compare model outputs against expected answers, reference data, or clearly defined scoring criteria.
For example, an eval for a contract review system might test whether the model identifies the correct clause, explains its significance, and cites the relevant passage. The results can reveal where the model performs well, where it fails, and what may need to improve before deployment.
In simple terms, LLM evaluation measures how well a model behaves. LLM training, on the other hand, is used to change that behavior.
Why is LLM evaluation important?
As LLMs take on more real-world tasks, teams need a reliable way to understand how well they perform and where they fall short. An LLM evaluation framework can help organize that process. From there, evaluation helps test whether a model is producing accurate, consistent, and useful results for its intended use.
Beyond measuring performance, evaluation can also help teams identify risks, compare models, guide improvements, and make more informed deployment decisions.
Here are 5 different ways LLM evaluation can help.
Model performance and reliability
A model producing one strong response doesn’t mean it will perform consistently. LLM evals test the model across representative inputs, edge cases, and repeated runs to see how reliably it completes the intended task.
For the contract review example, the system should be tested across different document lengths, clause wording, and levels of complexity rather than only on straightforward examples. This helps teams understand not just whether the model can succeed, but how consistently it does so.
Risk and failure identification
Evaluation can also reveal how a model fails and whether those failures matter.
A system may produce fluent responses while still omitting important information, making unsupported claims, exposing sensitive data, or taking an incorrect action. Some failures also carry more weight than others. Missing a termination clause in a contract, for example, is more consequential than producing an awkwardly worded explanation.
Testing for specific failure modes helps teams understand risks that may otherwise be hidden by an average performance score.
Consistent model comparison
LLM evaluation gives teams a structured way to compare models under the same conditions.
Models can be tested using the same tasks, inputs, tools, and scoring criteria so that differences in performance are easier to interpret. Teams can then compare quality alongside factors such as latency, cost, and reliability.
A lower-cost model, for instance, may not be the better choice if its outputs require significantly more human correction.
Model development and improvement
Evaluation results can help teams identify what needs to improve.
If a model repeatedly misses information that was available in the source material, the issue could be related to the prompt, retrieval system, available context, or the model itself. Teams can make changes and rerun the same evaluations to see whether performance improves.
Repeated evaluation can also catch regressions, where a change improves one part of the system but causes previously reliable behavior to deteriorate.
Confidence in deployment and decision-making
Evaluation provides evidence teams can use to decide how, or whether, to deploy an LLM system.
Results may support a broader rollout, a limited deployment, mandatory human review, or further development before launch. Those conclusions should remain tied to what was actually tested. Strong performance on contract extraction, for example, does not necessarily show that the same system is ready for every type of legal analysis.
For that reason, useful LLM evaluation starts with clearly defined tasks, realistic test cases, and measurable criteria for success.
What types of LLM evaluation can teams use?
Teams can evaluate LLM systems in different ways depending on when testing happens, how much of the system they assess, and what they need to learn from the results.
These approaches can be combined. A team might evaluate an entire application offline before deployment, then monitor individual components once the system is in production.
Offline vs. online evaluation
Offline and online evaluation differ mainly in when and where testing happens. Offline evaluation uses curated cases or replayed data under controlled conditions, while online evaluation examines production behavior and outcomes, sometimes incorporating delayed human feedback.
Offline testing can help teams compare systems or catch problems before deployment. Online evaluation can reveal how the system performs with real users, inputs, and operating conditions.
Offline doesn’t mean disconnected from the internet, and tests don’t always require reference answers. LangSmith’s evaluation concepts distinguish between these testing and monitoring uses. Online evaluation also shouldn’t expose users to untested, high-risk behavior.
Component-level vs. end-to-end evaluation
Component-level and end-to-end evaluations answer different questions about where a system succeeds or fails. Component checks test individual parts of the workflow, while end-to-end checks show whether the complete task succeeds.
A retrieval component might find the correct contract passage even though the final analysis misinterprets it. On the other hand, a useful final response could conceal a permissions violation along the way.
Testing both levels helps teams locate failures without assuming that strong individual components guarantee overall system success.
Model evaluation vs. application or product evaluation
Model evaluation tests how well a model performs under specific conditions, while application evaluation looks at the broader system people actually use. That can include prompts, retrieval, tools, permissions, interfaces, and the work users need to complete.
The distinction matters because a capable model doesn’t automatically make a reliable application. A model may perform well on an isolated task but still struggle when retrieval, tool use, permissions, or workflow requirements are added. Teams should therefore decide whether the evaluation is meant to measure the model itself or the complete product experience.
Once that scope is clear, teams can identify which parts of the application need to be evaluated and what evidence each one requires.
What elements of an LLM system should teams evaluate?
LLM evaluation should look beyond the model’s final response and examine the parts of the system that contribute to the result. Teams need to test prompts, retrieval, tools, routing, safety controls, and other components within representative user journeys to understand where the system performs well or breaks down.
Testing each part separately can show where the system performs well and where problems begin:
- Prompt performance: Test whether the system interprets instructions correctly, follows required output formats, and responds consistently across realistic variations in a request.
- Retrieval and generation quality: Retrieval-augmented generation (RAG) provides retrieved material as context for a model. Evaluate whether the system finds the right evidence and whether the final response uses that evidence accurately and completely.
- Agent behavior and tool use: Check tool selection, arguments, permissions, failure recovery, and observable outcomes rather than judging only the agent’s final response, as Mercor’s Agent Eval Systems explains.
- Fine-tuned model performance: Compare an adapted model with its baseline on held-out tasks to confirm intended improvements and detect regressions elsewhere.
- Routing and orchestration decisions: When evaluating model routing, test whether tasks reach the appropriate model, tool, or specialist step without sacrificing the required quality for lower cost.
- Safety and policy compliance: Test privacy handling, prohibited actions, appropriate refusals, and escalation against defined organizational policies. Passing these checks can show whether the system follows internal requirements, though it does not establish legal compliance.
Once teams know which parts of the system to test, they still need a reliable way to score the evidence those evaluations produce.
Popular LLM evaluation and scoring methods
There are several ways to evaluate and score LLM performance, and the right approach depends on what a team needs to measure. Some methods work well for objective tasks with clear right or wrong answers, while others are better suited to outputs that require human judgment.
The scoring method determines how an output is evaluated, while the metric captures the result. Choosing how to evaluate AI models means matching the scoring approach to the evidence a team needs rather than relying on a single score to represent every aspect of quality.
LLM-as-a-judge
LLM-as-a-judge evaluation uses one model to assess another model’s output. The model judge receives the task, candidate response, relevant evidence, rubric, and expected score format, then returns a score, pass/fail decision, or preference between responses.
This method of LLM evaluation can scale, but they also introduce their own risks. Zheng et al. identified biases related to answer position, response length, and the judge model’s own outputs. Before scaling, compare judge decisions with expert-labeled examples and retain the judge-model and prompt versions used. An LLM judge can support evaluation, but it shouldn’t be treated as independent ground truth.
Expert human evaluation
Expert human evaluation is useful when correctness depends on professional judgment that automated scoring may not capture well. Domain experts can assess accuracy, material omissions, practical usability, and whether an output meets professional standards.
In the contract example, a qualified lawyer should evaluate the substance of the clause analysis rather than simply checking whether a quotation appears. Human review also needs consistency controls. Reviewers should use shared rubrics, calibration examples, and a process for resolving disagreements. Mercor works with domain experts to define and evaluate realistic professional tasks where accuracy and context matter.
Hybrid evaluations
Hybrid evaluation combines automated checks alongside model judgement and expert review so that each method handles the criteria it’s best suited to assess. Deterministic checks, which return the same result for the same input, can confirm required fields or verify that a quotation appears in a contract. A calibrated judge can assess suitable rubric criteria, while experts can review ambiguous interpretations and consequential failures.
A strong review process should cover both high-risk cases and a representative sample of other outputs. Looking only at flagged results can miss errors the automated scorer failed to detect. LLM auto-evaluation works best as one part of the scoring process rather than as proof that every automated judgment is correct.
Whichever scoring method a team uses, it becomes more useful when it is part of a repeatable evaluation workflow.
What stages does the ideal workflow for LLM evaluation include?
A strong LLM evaluation workflow moves from defining the task to testing performance, analyzing failures, and monitoring the system over time. An evaluation framework organizes that process, while benchmarks provide standardized tests, scoring methods determine how judgments are made, and metrics capture the results. Each stage helps teams produce evidence they can use to improve the system or make deployment decisions.
In their practical guide to evaluating LLMs, Rudd et al. emphasize representative datasets, meaningful metrics, and robust execution as core parts of that process.
The following 8 stages put those principles into practice:
- Define the task and scope: Identify who will use the system, what context and tools it can access, what deliverable it should produce, and which failures are unacceptable. Separate minimum requirements from preferences.
- Build the evaluation dataset: Include representative tasks, difficult variations, and costly failure cases. Keep final test cases separate from examples used for tuning or development.
- Specify criteria and metrics: Define what each score measures and set acceptance thresholds before comparing models or systems. Identify any serious failure that should override a strong average score.
- Select and calibrate evaluators: Match code-based checks, model judges, and expert review to the criteria they are best suited to assess. Use shared examples to test scoring agreement and resolve unclear standards.
- Run controlled evaluations: Record the model, prompt, dataset, tools, and other important test conditions. Repeat suitable cases to measure variation instead of relying on the strongest run.
- Analyze results and failures: Look for recurring errors and differences across task categories. Review latency and cost alongside quality, and investigate cases where evaluators disagree.
- Apply the decision rules: Compare the results with the thresholds set earlier. Teams can then decide whether to deploy, limit use, improve the system, or gather more evidence.
- Monitor and refresh: Turn production failures into regression cases and rerun evaluations after meaningful changes to the model, prompts, tools, data, or operating conditions.
There is no universal passing score or required dataset size. The right level of coverage and uncertainty depends on the decision the evaluation needs to support.
The summary below shows the purpose of each stage:
| Evaluation workflow stage | Primary focus | Decision value |
|---|---|---|
| 1. Scope | Intended work | Establish boundaries |
| 2. Dataset | Representative cases | Test relevant conditions |
| 3. Criteria | Measurable standards | Fix the acceptance bar |
| 4. Evaluators | Scoring quality | Trust the assessment |
| 5. Execution | Comparable runs | Separate signal from noise |
| 6. Analysis | Failure patterns | Prioritize changes |
| 7. Decision | Predefined requirements | Deploy, limit, or defer |
| 8. Monitoring | Changing conditions | Detect deterioration |
Once the evaluation workflow is defined, teams need to choose metrics that show whether the system is meeting the performance criteria that matter for the task.
Which evaluation metrics are commonly used for LLM evaluation?
LLM evaluation metrics measure different aspects of model and system performance. The right metrics depend on the task and the questions an evaluation needs to answer.
Some metrics are reference-based. They compare outputs with expected labels, answers, or deliverables. Others are Reference-free metrics. These do not require an expected answer, although they may still use the prompt, source documents, policies, or a rubric. Some metrics can be used either way, while operational metrics such as latency and cost measure the performance of the system itself. The right metrics to track depend on the performance questions the evaluation needs to answer.
Reference-based metrics
Reference-based metrics are most useful when there is a reliable expected answer or label against which model output can be compared.
- Accuracy: Measures the proportion of predictions that match the correct answer for labeled tasks. Broad averages can still conceal poor performance on uncommon or high-risk cases.
- Precision, recall, and F1: Precision measures how many positive predictions are correct, while recall measures how many actual positives the system finds. F1 scores combine precision and recall using their harmonic mean.
- BLEU and ROUGE: These natural language processing metrics compare generated text with reference text. BLEU measures word-sequence overlap and includes a brevity penalty, while ROUGE measures overlapping text units through several variants.
Reference-free and criteria-based metrics
Reference-free metrics do not require a single expected answer. Instead, outputs can be assessed against the input, supporting evidence, a rubric, or other defined criteria.
- Relevance: Measures whether the response addresses the user’s actual request rather than only discussing the same topic.
- Groundedness: Checks whether claims are supported by supplied evidence, such as the contract passages cited in an analysis.
- Safety and policy adherence: Tracks defined violations and inappropriate refusals. Serious failures should remain visible instead of disappearing within an average score.
Metrics that can use either approach
Some dimensions can be evaluated against a known answer when one exists or through criteria-based scoring when it does not.
- Completeness: Measures whether the response includes all required information, including important clauses or qualifications that should not be omitted.
- Instruction following and task completion: Measures whether the system follows explicit constraints and whether it actually delivers the requested result. These are related but distinct aspects of performance.
- Robustness and consistency: Measures how performance changes across repeated runs, input variations, and difficult operating conditions.
System and operational metrics
- Latency: Measures how long it takes to produce a usable result, including retrieval, tool calls, and retries rather than only the initial model response.
- Cost: Measures the resources required to produce an acceptable result, including model usage, tools, retries, and necessary human correction.
No single metric can show whether an LLM system performs well across every dimension that matters. Teams should combine metrics with clear scoring criteria so they can see where performance is strong and where problems remain. In the contract example, finding the correct clause, quoting it accurately, and interpreting it correctly measure different parts of success.
Even well-chosen metrics can still produce misleading conclusions when the underlying data, scoring process, or test conditions are weak.
Limitations of LLM evaluation
LLM evaluation becomes less reliable when the data, scoring process, or test conditions don’t reflect the performance teams actually need to measure. A precise score can still answer the wrong question if the evaluation behind it is flawed.
Rudd et al. highlight the importance of dataset quality and execution conditions, while research on model judges shows that the scorer itself can introduce another layer of uncertainty. Several common weaknesses can affect how confidently teams interpret evaluation results:
- Unrepresentative data: A test dominated by simple contracts may miss failures in lengthy agreements, unusual wording, or incomplete source material. Results can look stronger than performance on the real range of work.
- Training or tuning contamination: Cases seen during training or repeatedly used for tuning may overstate how well the system performs on unfamiliar tasks.
- Weak reference answers: Incorrect, incomplete, or outdated expected results can reward mistakes and penalize valid alternatives.
- Ambiguous rubrics: Reviewers may apply vague criteria differently, making scoring differences look like changes in model performance.
- Judge bias: A model judge may favor persuasive presentation, response order, or other features that do not reflect the quality the rubric is intended to measure.
- Run-to-run variation: One successful execution cannot show how reliably the system behaves across repeated attempts. Repeated testing helps reveal whether performance is consistent.
- Changing conditions: Model updates, retrieval sources, tools, and production traffic can change how an earlier score should be interpreted. Results may need to be refreshed when important parts of the system change.
- Missed failures: Averages and limited samples may conceal rare but consequential errors or weak performance on particular task groups.
No evaluation can remove all uncertainty. These limitations should remain visible, and results should be interpreted in the context of what was actually measured. A stronger evaluation process can reduce uncertainty without suggesting that a single score provides complete evidence of reliability.
Best practices for evaluating LLMs
Strong LLM evaluation practices help teams produce results they can trust, compare, and use in real decisions. The most effective approach starts with realistic work, clear success criteria, and testing that reflects how the system will actually be used.
A stronger evaluation process starts with a few practical habits:
- Start with real work: Define the task, users, available context, and potential failure costs before choosing an evaluation tool or metric. Testing should reflect the work the system is expected to perform.
- Protect the final test: Keep held-out cases separate from prompt development and tuning so the evaluation measures performance on unfamiliar examples rather than rehearsed ones.
- Cover meaningful variation: Include representative work, difficult cases, and high-risk failures. Report important problem areas separately so they are not hidden by an overall average.
- Set decision rules early: Establish minimum requirements before comparing models or systems. Strong performance in one area should not compensate for an unacceptable failure elsewhere.
- Calibrate judgment: Use expert-reviewed examples to check whether evaluators apply the scoring criteria consistently. Resolve disagreements and refine unclear rubrics before relying on the results.
- Preserve the evidence: Track changes to prompts, models, datasets, rubrics, and tools, and retain the inputs and outputs behind each score. Clear records make results easier to review and reproduce.
- Repeat and monitor: Rerun evaluations after meaningful changes and turn production failures into new regression tests. Repeated testing shows whether improvements hold up over time.
Average scores rarely tell the whole story. Ask providers to show a failed case and explain what happened, what caused the failure, and how it affected the result.
The same principle applies to external benchmarks, which are only useful when teams understand what they actually test and how closely those tasks match the intended workflow.
What types of LLM evaluation benchmarks can be used to evaluate models?
LLM benchmarks test model performance on defined tasks so teams can compare results under consistent conditions. Some benchmarks measure broad knowledge or reasoning, while others focus on specific domains, professional tasks, or real-world workflows.
Metrics score benchmark results, evaluation harnesses run the tests, and leaderboards display comparisons. A strong benchmark score can provide useful evidence, but teams still need to confirm whether the model performs well on their own tasks and operating conditions.
General-purpose LLM benchmarks
General-purpose benchmarks test broad capabilities across a range of subjects and task types. Measuring Massive Multitask Language Understanding (MMLU), introduced by Hendrycks et al., assesses knowledge and problem-solving ability across academic and professional subjects using multiple-choice questions.
Its official repository supports broad capability comparisons, but MMLU doesn’t test tool use or whether a model can produce a complete professional deliverable. Choosing the correct answer to a contract law question is different from reviewing an unfamiliar agreement and producing a usable analysis.
Comparisons should distinguish the original MMLU benchmark from similarly named derivatives and use comparable testing and prompting conditions.
Domain-specific and real-world benchmarks
Domain-specific and real-world benchmarks test performance on narrower fields, tasks, or workflows. These evaluations can show whether a model can apply knowledge and reasoning in ways that more closely resemble the work teams expect it to perform.
Examples of specialized benchmarks include:
- LegalBench: This collaboration among legal and AI researchers covers classification, extraction, and rule application, largely in English and U.S. law. Scoring varies by task.
- SWE-bench: This benchmark evaluates software maintenance tasks by requiring code patches for repository issues. A patch succeeds when it resolves specified test failures without breaking tests that already pass. The result provides evidence about bounded software maintenance tasks rather than complete production readiness.
- APEX-Agents: Mercor’s expert-created environments cover investment banking, management consulting, and corporate law. Outputs and relevant artifacts or state changes are assessed against expert-authored binary rubric criteria. Partial-credit scores are reported separately from complete task success.
The most useful benchmark is one that reflects the kind of work a team needs the model to perform. For LLM reasoning evaluation, teams should judge observable outputs and task performance against defined standards rather than make assumptions about a model’s hidden reasoning.
Looking to evaluate LLMs with greater confidence?
Build your evaluation around realistic work and standards that qualified reviewers can defend. See how Mercor combines domain expertise, task design, and repeatable assessment for your workflows.
Explore enterprise evalsFrequently Asked Questions
What are the best tools for LLM evaluation?+−
The best LLM evaluation tool depends on what a team needs to test and how results will be scored. Teams should also consider the types of environments, benchmarks, and grading methods a tool supports. Harbor supports agent and language model evaluation across benchmarks and environments, while Archipelago supports evaluating AI agents in configurable environments with built-in grading.
What are rubrics in LLM evaluation?+−
Rubrics define the criteria and scoring rules used to evaluate model outputs. A contract-review rubric could assess whether the system identifies the relevant clause, cites the supporting passage, and flags a material omission. The criteria define what matters, while thresholds determine what passes. The same rubric can guide both human and model-based reviewers.
How do you write evals for LLMs?+−
Start by defining the inputs, permitted context and tools, expected results or rubrics, scorers, and pass/fail rules. Code can check whether a quoted clause appears in the contract, while an expert assesses whether the model interpreted it correctly. Teams can retain the evaluation case for later regression testing to see whether future changes improve or weaken performance.
