Aug 27, 2026Enterprise

Your AI should earn underwriting authority the way a new underwriter does

Ashley Liu, and Vaithiyanathan Sundaram

With Brandon Lei, Braden Holstege, Joe Kaplan, Alex Gonzalez, and Imani Parker-Fong

In one paragraph

Your submission tool went live and cycle time on the small commercial book improved. Twelve months later, it is performing much as it did at launch.

Underwriters improve through regular file reviews and earn greater authority as their judgment proves reliable, but most AI systems have no equivalent process. Insurance already has a model for managing this risk. Some underwriting decisions can be made independently within defined limits, while others require referral or specialist review. AI touches 17% of underwriting tasks today, and executives expect that to reach 75% within three years.¹

As AI takes on more of the underwriting workflow, carriers can apply the same model by turning existing file reviews into continuous evaluations and using the results to determine what AI can handle independently, what requires review and where an underwriter still needs to make the decision.

1. AI needs the equivalent of a file review

The challenge is that underwriting quality can be difficult to evaluate quickly. A case file or quote letter is rarely simply right or wrong, and much of the underlying submission data still arrives in documents.

In some cases, the ultimate test may not come until a claim arrives. By then, the model, data feed and underwriting guidelines may all have changed, making it difficult to trace the original decision back to what the system saw and produced at the time.

Reviewing every file manually is not a practical answer either. The underwriters best equipped to assess the work are already constrained for time. Underwriters spend 41 to 43% of their time on administrative work,² and 64% of underwriting executives cite talent as one of the external forces they expect to have the greatest impact on their business.¹

2. Turn the file review into an eval

Most carriers already use peer review or underwriting audits to assess the quality of underwriting decisions. A senior underwriter reviews a sample of bound files against the carrier’s guidelines, checking whether the exposure was captured correctly, the classification was appropriate, the loss history was reconciled, the right referrals were made and the rate was justified. The limitation is that these reviews typically happen after the fact and cover only a fraction of the book.

For AI, carriers can take the same criteria and turn them into a scorecard that evaluates every file the system produces before a quote reaches the broker. In technical terms, this is an eval.³ Mercor uses this approach to evaluate AI on complex professional work, and in underwriting the evaluation can be built around the carrier’s own files, guidelines and standards.

3. Build the eval from real underwriting work

An eval has five components that together define the task, what good work looks like and how the AI’s performance is measured.³

The file. A sandboxed copy of what the AI sees including the broker thread, the statement of values, five years of loss runs, survey and guideline in force.

The task. One defined deliverable at a time so the submission-to-quote process can be evaluated as distinct pieces of work.

The benchmark. You do not need to start from scratch. Real files that experienced underwriters have already reviewed provide the benchmark, along with the source material and context needed to establish what good work looks like.

The scorecard. Weighted criteria with partial credit that preserve the complexity of a real submission. If the valuation on location five is stale, the AI needs to flag it.

The marker. The automated reviewer that applies the scorecard, calibrated against expert judgment until it reliably reflects the underwriting team’s quality bar.

Benchmarks also need to reflect the context in which an underwriting decision was made. A file written in a hard market can penalize a correct decision in a softening one, so each benchmark carries an effective date, guideline version and rate level, and is updated as those conditions change.³

4. Authority should follow performance

Once the quality of the AI’s work can be measured, carriers can use that evidence to determine where the system operates independently and where an underwriter remains in the loop. The existing tiers of underwriting authority provide the structure.⁶

Straight through. The AI can act within agreed limits when its work clears the quality threshold, the risk falls within defined appetite and the underlying data is reliable. The system should also preserve a clear record of the information and logic behind each decision so the work can be audited and explained.

Prepare and refer. The AI assembles the file and recommends an action, but an underwriter makes the decision.

Specialist review. Complex risks, incomplete data and work below the quality threshold stay with a specialist.

The practical place to start is with work an underwriter can verify quickly. Zurich reports up to two hours saved per submission on U.S. Middle Market underwriting narratives, and a Sixfold case study of Zurich North America reports that 80% of submissions were processed in the initial phase while the underwriter retained the risk decision.⁴ Accenture reports that QBE can process 100% of broker submissions in the product lines where its AI underwriting solutions are in production.⁷

As the carrier collects more evidence, the level of authority can change. If the AI consistently clears the required quality threshold for a particular class of risk, its authority can expand within defined limits. If performance declines after a model, guideline or data change, that authority can contract. If a class continues to perform poorly, the referral stays.

The referral path also preserves an important part of underwriter development. Junior underwriters continue to see the facts, the recommendation and the expert judgment together rather than having the work that builds judgment disappear entirely into automation. Accenture’s survey points in the same direction, with 81% of respondents expecting AI to create new roles and 76% expecting easier knowledge transfer.¹


Diagram showing how AI underwriting files are routed to straight-through processing, underwriter referral, or specialist review based on evaluation results.

Figure 1. How each file is routed. The AI case manager assembles and prices the file, and each file then lands in one of three levels of authority. The automated file review runs before and after the quote, so underwriter overrides and claims outcomes feed back into the scorecard. Structure adapted from note 6.


AI underwriting workflow from submission receipt through evaluation, referral, underwriting decision, and quote.

Figure 2. One submission, end to end. The checks between steps are automated markings, so a file below the threshold is held, reworked or referred before a broker sees it. Process adapted from note 6.

5. Use overrides and outcomes to improve the eval

Every reviewed file creates a record of what the system saw, what it produced, how it scored and where an underwriter disagreed. Repeated overrides can reveal gaps in the scorecard, while claims outcomes can show which criteria actually matter to underwriting performance.³

The same feedback loop also makes changes to the system easier to manage. If performance drops after a new model, data source or guideline is introduced, the carrier can see where quality changed and adjust the authority accordingly.

The system should also preserve a clear record of the information and logic behind each decision so the work can be audited and explained.

6. Start with one book

The first ninety days do not need a new system. Pick one book where submissions are frequent. Collect twenty files experienced underwriters have already reviewed and work with a small number of them to turn their judgment into an explicit scorecard.³

Waiting for better technology is not a plan. On Mercor’s open benchmark of expert built professional work, the leading model passes 43.5 % of tasks at one attempt, so even the strongest systems still fail a majority of them.⁵ Carriers already know how to manage high stakes judgment without requiring perfection. They set limits, review the work and expand authority when the evidence supports it.

AI should earn its authority the same way.

Notes

  1. Accenture, Underwriting Rewritten, August 2025. Survey of 430 senior underwriting executives across Life, Commercial P&C and Personal P&C in 11 countries, with 25 follow-up interviews. Source of the 17% and 75% AI task figures, the finding that 64% of executives identify talent as an external force expected to have a major impact, and the findings that 81% expect AI to create new roles and 76% expect easier knowledge transfer.
  2. Capgemini, Unleashing Growth: The Evolving Role of Underwriters, 2024. Reports that commercial and personal lines underwriters spend 41% to 43% of their time on administrative activities.
  3. Mercor, Agent Eval Systems and Enterprise Evals, 2026. Source for Mercor’s approach to defining tasks, benchmarks, scorecards and automated reviewers for evaluating AI work. Evaluations can be used to grade live traffic, gate delivery or test changes before release. The build sequence, calibration process and use of expert-defined quality criteria described in this article reflect Mercor’s methodology.
  4. Zurich Insurance Group, Annual Report 2025, and Sixfold’s Zurich North America case study. Zurich reports up to two hours saved per submission on U.S. Middle Market underwriting narratives. Sixfold reports that 80% of submissions were processed in the initial phase, with the underwriter retaining the risk decision. Mercor was not involved in this work.
  5. Mercor, APEX Agents, 2026. An open benchmark of 480 rubric-graded tasks across 33 expert-built environments in investment banking, management consulting and corporate law. As of August 2026, the leading entry passes 43.5% of tasks at one attempt. Rankings may change as new models are released, so the current result should be checked before publication.
  6. McKinsey, The Future Underwriting Operating System: From Inbox to AI Nerve Center, June 2026. Source for the human-governed underwriting operating model and the three routes described here as straight through, prepare and refer, and specialist review.
  7. Accenture, 2025 client example. QBE Insurance Group is scaling AI-powered underwriting solutions co-developed with Accenture. In the product lines where those solutions are in production, Accenture reports that QBE can process 100% of broker submissions. Mercor was not involved in this work.