Large language model (LLM) evaluation metrics are useful when they measure the work you need done, not just its performance on a convenient test. The model with the highest overall score may still miss a contract obligation that matters to your business.
The right evaluation therefore depends on the task, the evidence available to judge it, and the failures a team cannot afford.
This guide explains what LLM evaluation metrics measure, the main types of metrics, how they are scored and interpreted, and how to choose the right combination for a real-world workflow.
What are LLM evaluation metrics?
LLM evaluation metrics measure specific aspects of how well a model or AI application performs. A metric might assess whether an answer is correct, whether its claims are supported by source documents, whether an agent completed a task, or how long the system took to produce a usable result.
It helps to distinguish a metric from the broader evaluation process:
- A metric defines what is measured.
- An evaluation applies one or more metrics to a set of test cases.
- A benchmark provides a standardized set of tasks and a scoring procedure.
- A grader or scorer determines how an individual result is assessed.
Those distinctions matter because evaluating a base model is different from evaluating an application built around one. An application may add retrieval, tools, instructions, memory, or other components that introduce additional ways for the system to succeed or fail.
For that reason, LLM performance should not be judged from a leaderboard score alone. OpenAI's evaluation guidance begins with an evaluation objective and representative test data, while Anthropic's guidance emphasizes tasks, graders, and the evaluation harness used to run them.
Why do LLM evaluation metrics matter?
LLMs can produce convincing outputs without performing reliably enough for the work they are expected to do. Evaluation metrics give teams a consistent way to measure that performance and understand where problems remain.
A repeatable scorecard provides evidence teams can use to:
- Measure model quality consistently: Apply the same definition of a correct or acceptable result to each case.
- Compare models or versions using the same criteria: Keep inputs, evidence, tool access, and grading rules consistent so score changes are easy to interpret.
- Identify strengths and common failures: Break down results by domain, difficulty, and failure type instead of relying on a single average.
- Track performance changes over time: Rerun a stable test set after changes to prompts, retrieval, tools, or models to catch regressions.
- Support decisions about deployment and improvement: Use measured trade-offs and critical-failure criteria to decide what to fix and whether a system is ready for release.
A high score doesn't eliminate risk; it reflects performance on the cases and conditions that were tested. Results may also vary between attempts, which is why repeated trials and clearly defined baselines can matter when evaluating nondeterministic systems.
What are the main types of LLM evaluation metrics?
LLM evaluation metrics measure different aspects of performance. A contract-review assistant, for instance, might need to retrieve the right information, produce an accurate answer, complete a task, and do so at an acceptable cost. The following metric types can overlap, and teams can score them using rules, human reviewers, or another model.
Output quality metrics
Output quality metrics assess whether a response meets the task's requirements.
- Correctness measures whether an answer is accurate
- Relevance checks whether it addresses the request
- Completeness identifies missing information
- Instruction adherence measures whether the model follows directions
- Groundedness evaluates whether claims are supported by the supplied evidence, but that support doesn't guarantee factual correctness. A response that sounds convincing can still be wrong.
- Safety: Whether the response follows defined safety or policy requirements.
These dimensions should not be treated as interchangeable. A response can be relevant but incorrect, or grounded in a source that is itself inaccurate. Likewise, a fluent and persuasive answer does not establish that the underlying information is correct.
RAG evaluation metrics
Retrieval-augmented generation (RAG) gives an LLM external information to use when answering questions. RAG evaluation metrics assess both what information was retrieved and what the model did with it.
Common RAG metrics include:
- Retrieval precision measures how much of the retrieved material is relevant, while context recall assesses whether the system retrieved the required information against a reference.
- Context recall: Measures whether the retrieval system found the information required to answer the question, when an appropriate reference is available.
- Faithfulness measures whether the answer's claims are supported by the retrieved material.
Ragas context precision also considers where relevant material appears in the retrieved results. These metrics help distinguish between different failures. A system might miss an important contract clause during retrieval or retrieve it but generate a claim the source doesn't support.
Agent and task performance metrics
Agent evaluations focus on whether the system successfully completes the work, not only on whether its final response sounds correct.
Metrics may assess whether an agent:
- Completed the requested task.
- Selected the appropriate tools.
- Used those tools correctly.
- Reached the expected final state.
- Followed required policies or approval rules.
- Recovered appropriately from errors.
That is why evaluating AI agents often requires examining both the final result and the actions that produced it.
Repeated trials can also provide information about reliability. pass@k estimates the probability of obtaining at least one successful result across k attempts, while pass^k estimates the probability that all k attempts succeed. Teams should define the number of trials and permitted retries consistently when comparing systems.
Multi-turn and conversational metrics
A single good answer does not guarantee a successful conversation.
Multi-turn metrics can assess whether an LLM remains consistent across an interaction, retains relevant user requirements, asks for clarification when appropriate, and ultimately resolves the user's request.
Evaluations may therefore consider:
- Requirement retention.
- Consistency across turns.
- Appropriate clarification.
- Error recovery.
- Conversation-level task completion.
When simulated users are used in testing, their behavior should reflect realistic interactions rather than artificially making the system easier or harder to evaluate.
Traditional NLP metrics
Traditional natural language processing (NLP) metrics remain useful for specific tasks.
Accuracy, precision, recall, and F1 help assess classification or information extraction using defined labels.
Other metrics compare generated text with reference text. BLEU measures overlap with reference translations, while ROUGE measures overlap with reference summaries. BERTScore assesses contextual similarity between words and phrases.
These metrics provide useful signals, but similarity does not establish factual correctness.
Likewise, perplexity measures how well a model predicts text, but that doesn't mean it can complete a real-world task successfully.
System and operational metrics
An LLM system also needs to perform efficiently enough for its intended use.
Operational metrics can include:
- Time to first token: How quickly a response begins.
- End-to-end latency: How long it takes to produce the complete usable result.
- Throughput: How much work the system can process under load.
- Error and timeout rates: How often workflows fail operationally.
- Cost per successful task: The total cost of successful and failed attempts, retries, tool use, and other resources divided by completed tasks.
These measures should be interpreted alongside quality. A faster or cheaper system is not necessarily better if it completes fewer tasks correctly or requires substantially more human correction.
What is the difference between reference-based and reference-free evaluation?
The difference between reference-based and reference-free evaluation lies in the evidence used for comparison, not whether the grader is human or automated.
| Approach | Required evidence and examples | Useful when | Boundary |
|---|---|---|---|
| Reference-based | A target answer, gold label, relevant-document set, or expected final state, such as clause recall against expert-labeled obligations | You can compare an output against an independently established target | An incomplete or incorrect reference can penalize valid alternatives or reward omissions |
| Reference-free | No target answer; evaluation uses the task, retrieved passages, constraints, or an anchored rubric, such as assessing answer faithfulness against supplied passages | Acceptable answers vary but supporting evidence or specified behavior can still be assessed | Support doesn't prove that the source is correct or that all required information was retrieved |
Some retrieval-recall methods require reference contexts or claims from a reference answer. In contrast, Ragas's faithfulness metric uses retrieved context without requiring a reference answer. Specify which implementation is being used, and review the quality of any references before relying on the results.
How are LLM evaluation metrics scored?
LLM evaluation metrics can be scored in several ways, depending on what needs to be measured and the available evidence. A criterion defines success, such as identifying every critical contract obligation, while the scoring method determines how to check whether that criterion was met.
Rule-based and deterministic scoring
Rule-based scoring works best when the expected result is clearly defined. Teams can use exact or normalized matches, validate output formats, run executable tests, or verify whether a system reached the correct final state.
An evaluation record should capture the case ID, input, output, any reference answer or retrieved evidence, relevant tool activity or final state, and latency or cost. The scorer then returns a pass or fail outcome with a reason for the result.
Teams should define permitted normalization and numerical tolerances in advance. Passing a format or schema check confirms that an output follows the required structure, not that its content is correct.
Statistical and similarity-based scoring
Statistical metrics turn defined outcomes into numerical scores. For tasks with positive and negative labels, precision = TP/(TP + FP) measures how often predicted positives are correct, while recall = TP/(TP + FN) measures how many actual positives were identified. F1 = 2PR/(P + R) balances precision (P) and recall (R). Here, TP means true positives, FP means false positives, and FN means false negatives.
Teams must define these outcomes for their task. In contract extraction, for example, the unit might be an individual obligation rather than an entire document. Teams should also specify matching rules and how to handle zero denominators.
Text-similarity metrics work differently. BLEU measures reference-text overlap using modified n-gram precision and a brevity penalty, ROUGE measures specified types of overlap, and BERTScore assesses contextual similarity. These measures indicate how closely outputs resemble references, but they don't establish factual correctness.
Retrieval metrics also need precise definitions. Simple document precision@k divides the number of relevant documents among the top k retrieved by k using predefined relevance labels, while Ragas's context precision can also account for where relevant results appear in the ranking.
Model-based and LLM-as-a-judge scoring
LLM-as-a-judge scoring uses another model to evaluate an output against defined criteria. The judge receives the task, candidate response, rubric, and any evidence or reference material needed for the assessment. For example, Ragas's faithfulness metric calculates a score by dividing the number of generated claims supported by the retrieved material by the total number of generated claims.
Results should include the score and supporting evidence, with an option such as "insufficient evidence" when the judge cannot make a reliable decision. Teams should also define how to handle responses with no claims or cases where grading fails.
Keep the judge model, prompt, and rubric versions fixed when comparing results. For pairwise comparisons, swapping the answer order and re-running the judge can help reveal position bias. Teams should also check whether judges favor longer responses and validate automated scores against expert labels. Mercor's guide to how LLM-as-a-judge works explores these considerations further.
A separate metric may still be needed to assess qualities the judge doesn't measure. For example, an answer may be well grounded in its sources but omit important information.
Human and expert evaluation
Human evaluation is useful when assessing output quality requires professional judgment. Domain experts can review blinded cases against a defined rubric, compare their ratings, and resolve disagreements.
Teams can measure reviewer consistency using statistics such as Cohen's kappa, which assesses agreement between two categorical raters beyond chance. However, it measures consistency, not whether their judgments are correct.
Expert review should cover representative cases and difficult edge cases to determine whether automated scoring reflects professional standards. Expertise doesn't eliminate subjectivity, and reviewer agreement alone doesn't prove correctness.
How to interpret evaluation metrics?
To interpret evaluation metrics, start by checking each score's definition, unit, denominator, scale, and aggregation method. A reported value of 90, for example, could represent 90% of obligations identified, a point on a grading scale, or 90 milliseconds of latency. Check whether the calculation also accounts for timeouts and abstentions.
Aggregation matters, too. A macro average gives each defined group, such as legal, medical, or finance cases, equal weight. A pooled or micro measure gives greater weight to groups with more cases. Both can conceal serious failures, so teams should examine individual risk areas alongside overall results.
Compare systems using the same held-out cases, evidence, rubric, and operating conditions. Small score differences aren't automatically meaningful, particularly when results vary across repeated trials. When multiple observations come from the same case or conversation, estimate uncertainty at that level rather than treating them as independent results. Use a sampling method appropriate to the evaluation design.
A compact evaluation report should include:
- Evaluated cases, exclusions, sample size, and an appropriate confidence interval
- Model, prompt, retrieval, tool, grader, and rubric versions
- Results by domain and difficulty, including failed-run and abstention rates
- Quality gates alongside latency and cost under the same workload
For the hypothetical contract-review assistant, a higher average answer-quality score cannot compensate for missed critical obligations. Examine these failures under consistent acceptance rules before choosing a lower-cost system. Avoid comparing raw scores from unrelated datasets or applying a universal deployment threshold.
Perplexity comparisons also depend on how text is tokenized and how the context window is handled. Anthropic emphasizes repeated trials and valid graders when interpreting evaluation results.
How to choose the right LLM evaluation metrics?
Start with the decision the evaluation needs to support, then choose metrics that reveal the most important failures. OpenAI's evaluation design sequence begins with the objective and test data rather than a preferred metric.
Start with the task and desired outcome
Define what the system needs to accomplish and what success looks like. In our hypothetical contract-review workflow, each test case involves one contract package. Success means identifying the required obligations and supporting them with relevant passages, not simply producing fluent prose. Also identify who will make the final release decision.
Identify the highest-risk failure modes
Consider which failures would have the greatest impact and how likely they are to occur. Missed critical obligations, clinical omissions, incorrect financial calculations, misleading consulting recommendations, and software regressions all require different checks.
Treat unacceptable failures as hard gates rather than allowing stronger scores elsewhere to offset them. Other quality measures can still be compared on a scale. These examples illustrate evaluation design, not legal, clinical, or financial advice.
Determine what evaluation data is available
Choose metrics based on the evidence available. Expert-approved answers or labels support match-based scoring, while relevance judgments support retrieval recall. Supplied passages allow faithfulness checks, and tool traces or final states help evaluate agents.
Use representative data that your team has permission to evaluate, de-identifying sensitive information where required. Reserve cases for held-out testing, and obtain expert labels where existing evidence is insufficient.
Decide whether the workflow is single-turn or multi-step
Match the evaluation to the workflow. Scoring the final answer may be sufficient for a single-turn task. If the system retrieves information, uses tools, or holds conversations, examine both the outcome and the steps leading to it.
These checks may include tool inputs, changes in system state, error recovery, and handoffs. A correct individual step doesn't guarantee that the entire workflow succeeded.
Choose a mix of complementary metrics
For contract review, combine expert-labeled recall of critical obligations with groundedness, instruction adherence, verified task completion, latency, and cost. Each metric captures a different quality requirement or operational constraint. Keep critical-failure gates separate from broader quality scores so strong performance elsewhere cannot offset an unacceptable failure.
Validate metrics against real-world performance
Compare metric scores with expert judgments on representative cases, examining false passes and false failures, especially in high-risk tasks. Then validate the metrics on held-out workflows to assess their reliability beyond the cases used to develop them.
Reevaluate the scorecard after meaningful changes to the model, prompts, data, or judge, and establish a process for resolving disagreements with expert reviewers. Compare the results with actual operational outcomes. User satisfaction offers useful feedback but cannot establish correctness in high-stakes work.
Mercor's guidance on evaluating AI models for your company outlines the broader selection process.
What are common mistakes when using LLM evaluation metrics?
LLM evaluation metrics can be misleading when teams measure the wrong things, compare results under different conditions, or rely too heavily on a single score. The table below outlines common mistakes and corrections.
| Mistake | Correction |
|---|---|
| Treating a useful signal as proof of correctness. | Overlap, similarity, fluency, and source support measure different aspects of performance. Check factual accuracy and required coverage separately. |
| Optimizing for the test rather than the task. | Reserve representative cases for evaluation, and update them as the workflow evolves. A public LLM benchmark cannot guarantee performance in a private contract-review workflow. |
| Comparing results under different conditions. | Use consistent datasets, configurations, reference materials, scoring rules, and graders. Document any changes before interpreting performance trends. |
| Relying on a single overall score. | Examine high-risk tasks separately, and apply critical-failure gates to prevent strong averages from concealing serious errors. |
| Trusting unvalidated graders or excluding failed cases. | Compare automated judgments with expert labels, and establish how timeouts, abstentions, and grading failures are counted. |
OpenAI's evaluation guidance makes a similar point: evals should reflect how the system is used in the real world.
When should you create custom LLM evaluation metrics?
Create custom LLM evaluation metrics when standard measures fail to capture an important requirement of the task. For example, a contract-review assistant could use a custom coverage metric to assess whether it identifies obligations that qualified reviewers consider critical.
A clear metric specification should include:
- What the metric evaluates: Use one contract package, the assistant's extracted obligations and cited passages, and an expert-approved list of critical obligations.
- How matches are scored: Qualified reviewers define what constitutes a valid match. Calculate coverage by dividing correctly matched critical obligations by all labeled critical obligations. Count each obligation once, and require supporting passages.
- How unusual cases are handled: If a contract has no critical obligations, mark its coverage score as not applicable and report it separately. Record missing outputs, unsupported additions, and grading errors rather than excluding them.
- How results affect decisions: Report coverage for each eligible contract and the proportion of eligible cases meeting the team's defined threshold. A system passes only if it meets the required quality standards without triggering a critical-failure rule. No single threshold suits every workflow.
- Who owns the standard: Qualified reviewers determine which obligations are critical and resolve ambiguous matches. Record the metric definition, rubric, and evidence versions to maintain consistency.
Before relying on a custom metric, test it against clear successes, failures, and borderline cases. Validate it using held-out contracts, and reassess it as the workflow changes. Mercor's enterprise evaluations use expert-built criteria to assess real workflows, although evaluation scores alone cannot establish regulatory compliance or professional safety.
Evaluate your LLM performance against real professional standards
The right mix of metrics is only part of building a useful evaluation. Mercor Enterprise Evals helps teams build expert-informed evaluation tasks and rubrics around real professional workflows.
Explore enterprise evals