Will AI replace doctors? What do the benchmarks say?

Will AI replace doctors? What do the benchmarks say?

Current frontier AI models can already perform many physician tasks well, but replacing a doctor requires far more than excelling on a single test or capability score.

A doctor’s work includes hundreds of distinct tasks, from reading a chest X-ray to calming a scared patient in the middle of the night. AI models perform unevenly across this spectrum. Some tasks they handle at or near physician level, while others remain beyond their capabilities.

Much of what gets published on this topic is speculation rather than detailed analysis. While one source may say that AI will replace doctors within a decade, another will claim it never will. However, these claims lack rigorous testing.

What's missing is a direct assessment of how today's AI models perform on the actual work physicians do, graded against real clinical standards rather than vendor demonstrations or licensing exams.

Benchmarks such as Mercor’s AI Productivity Index (APEX) are designed to fill this gap. They evaluate how frontier models perform on physician tasks, scoring performance task by task.

What does Mercor's APEX benchmark for general practitioners measure?

Mercor’s APEX tests whether frontier AI models can perform economically valuable professional work, including the work of a primary care physician.

The general practitioner track was created with physicians affiliated with the University of Pennsylvania, Northwestern, Cornell, Brigham and Women's, and Mount Sinai, with guidance from cardiologist Eric Topol, founder of the Scripps Research Translational Institute.

How APEX measures physician work

Instead of relying on multiple-choice questions, APEX presents models with realistic clinical cases paired with source documents such as patient charts. It then grades the models' responses against rubrics written by experts.

The benchmark includes 100 cases for each profession it evaluates across medicine, law, finance, and consulting. These cases are designed to test whether models can complete valuable professional tasks rather than simply recalling information they may have encountered during training.

What types of medical tasks are evaluated?

The medical cases reflect the type of work found in a general practitioner’s caseload. For example, one sample task gives the model a patient’s records and asks whether their prescription dose should change given their stable liver test results. It resembles the kind of medication-management judgment a physician might routinely make.

Other cases cover diagnosis, test interpretation, and treatment planning. Models must review the supplied information, identify the details that matter, and produce a reasoned response rather than select an answer from a list.

This distinction matters because APEX doesn’t only measure whether a model can recall medical facts. It measures whether the model can apply those facts to realistic clinical work, use the available records, and produce a response that meets standards defined by medical experts.

How well do frontier AI models perform on physician work?

On the current APEX General Practitioner leaderboard, the top model scores 71.6%, with most frontier models clustered between 58% and 67%.

These results demonstrate that frontier AI models already display substantial capability on realistic physician tasks. However, a high score still leaves too much room for error to replace specific roles.

What does the benchmark score mean for the future of doctors?

The evidence points to AI augmenting doctors' roles, not replacing them. Frontier models can already assist with a meaningful share of a physician's work.

However, for such models to be fully automated, they would need to demonstrate far more consistent performance and the ability to handle incomplete information, physical examinations, and unpredictable patient interactions.

Which physician tasks do frontier models perform well today?

APEX doesn’t publicly break down model performance by individual task category, so its leaderboard can’t show exactly which types of physician work account for the strongest scores.

Still, the benchmark demonstrates how frontier models can apply supplied medical information to realistic clinical tasks and produce responses graded against expert standards.

Broader medical research suggests that AI models perform best when the task:

  • Starts with complete patient records
  • Has a defined clinical question
  • Requires reasoning over supplied information
  • Produces a written recommendation
  • Can be reviewed by a physician

Examples of these tasks include:

  • Reviewing patient charts
  • Performing medication calculations
  • Identifying relevant medical history
  • Drafting clinical recommendations
  • Summarizing findings

What do frontier models still struggle to do?

The gap between benchmark scores and the capability needed for independent practice doesn't stem from one missing feature. Instead, it's a cluster of limitations that include the ability to:

  • Perform physical examinations or any kind of hands-on assessment
  • Handle edge cases that don't resemble training data
  • Take legal and moral accountability for clinical outcomes
  • Gather information from incomplete or unreliable patient accounts

Models generally face greater difficulty when they must synthesize incomplete records, sequence multistep workflows, or decide what additional information to gather. These kinds of open-ended judgment calls make up much of a physician’s day.

Measuring how AI is benchmarked across professional work with expert-graded rubrics reveals points of failure in AI model training that vendors can use to improve future releases..

A model may show strong performance on a narrow range of tasks, but that's not the same thing as practicing medicine. Productivity benchmarks like APEX evaluate AI on the actual work.

Which medical specialties are changing fastest?

The clinical fields where AI is making changes the fastest are those where a large share of the work involves image interpretation, structured data, or repetitive cognitive tasks. These are the categories where models already perform well.

The table below summarizes the medical specialties where AI exposure is currently highest, which tasks are increasingly being handled by AI, and which aspects remain dependent on human expertise:

SpecialtyAI exposureWhat AI increasingly handlesWhat stays human
RadiologyHighScreening reads, triage, quantificationComplex cases, integration with clinical context, final reads
Pathology and dermatologyHighImage classification, lesion flaggingBiopsy decisions, atypical presentations, patient counseling
Primary careModerateDocumentation, inbox triage, care-gap alertsExamination, longitudinal relationships, undifferentiated symptoms
Surgery and proceduresLow to moderatePre-op planning, imaging guidanceThe procedure itself, intraoperative judgment
Psychiatry and emergency medicineLowNote drafting, risk-score supportRapport, de-escalation, real-time triage where there's uncertainty

Image-based specialties show the strongest benchmark performance

Radiology, pathology, dermatology, and ophthalmology consistently show the highest AI exposure because many of their core tasks involve matching visual patterns to a known category, which is exactly the type of task that current models perform well.

That doesn't mean that radiologists are becoming obsolete. Rather, it means that a growing share of first-pass reads gets triaged or pre-screened by a model, with a physician confirming, correcting, or overruling as needed.

Primary care benefits from workflow automation

Primary care sees moderate exposure to AI, and it's primarily concentrated on administrative tasks, such as documentation, coding, after-visit summaries, and message triage. The core of a general practitioner's work, interpreting symptoms in the context of a patient's history, remains a fundamentally human task based on current benchmarks.

Procedural specialties remain difficult to automate

Surgery and other hands-on fields show the lowest exposure to AI because benchmark performance on cognitive tasks says nothing about a model's ability to hold a scalpel or adapt mid-procedure to unexpected events. AI's role in these fields remains purely pre- and post-procedural.

Human interaction keeps psychiatry and emergency medicine physician-led

Psychiatry and emergency medicine resist AI automation for a different reason: much of the work depends on real-time human judgment under emotional or physical crisis, a context no current benchmark task actually replicates.

A model can draft a psychiatric intake note, but it can't sit with a patient in crisis and decide, in the moment, whether they're safe to go home.

What does broader medical research say about the future of doctors?

APEX demonstrates how frontier AI models perform on realistic physician tasks. Broader medical research helps explain where those strengths come from and where they still remain limited.

Across studies, AI performs best on tasks that involve structured knowledge, documentation, and image-based interpretation, while physical examinations and open-ended clinical judgments remain far harder to replicate.

Medical knowledge and clinical reasoning

AI models perform strongly on medical knowledge retrieval and stepwise reasoning tasks, often reaching or exceeding the performance of an average physician on recall-heavy questions.

That reflects the strength of the underlying training data, which contains enormous volumes of medical text, more than any deep clinical insight. Knowing the right answer to a well-posed question is not the same skill as knowing which question to ask a confused, anxious patient in front of you.

Documentation and administration

Drafting clinical notes, summarizing patient histories, and generating discharge instructions are among the most reliably strong applications for frontier models because inputs are structured and output formats are predictable.

This is also the area where physicians report the most immediate benefit from AI, since documentation is widely cited as a leading driver of burnout. AI can reduce administrative work without replacing clinical decision-making.

Diagnostic interpretation

When inputs are standardized images, AI performance is often strongest. The MASAI randomized trial in breast cancer screening found that AI-supported mammography screening increased cancer detection and reduced radiologists' screen-reading workload. Overall, models perform best on tasks that:

  • Rely on written or imaging inputs rather than physical presence
  • Have well-defined correct answers a rubric can grade
  • Repeat at high volume with limited case-by-case ambiguity
  • Draw on published medical knowledge rather than patient-specific context

Physical examination remains outside benchmark performance

Few benchmarks can measure a model's ability to, for example, palpate an abdomen, listen for a subtle heart murmur, or notice the guarded way a patient moves.

That's not a temporary gap in the training data. It's a structural limitation of a system that works with text, images, and structured inputs rather than physically interacting with a patient.

Clinical judgment extends beyond pattern recognition

AI models often arrive at correct conclusions through pattern matching rather than the kind of reasoning a physician would use. While that distinction may not matter when the case is textbook, it becomes apparent when the case is less straightforward.

It's unlikely that doctors will be replaced by AI. Instead, the evidence suggests that the physician role will shift toward supervising, verifying, and integrating with AI. As a result, doctors’ task lists will look different, with less typing, more oversight, and more time spent on the work that only humans can do.

Furthermore, as frontier models grow more capable, clinical expertise becomes more valuable for building them, creating a growing range of new AI opportunities for physicians emerging across the industry.

Clinicians now write, grade, and stress-test the evaluations that determine whether medical AI is safe to deploy. Physicians help evaluate AI by contributing their clinical judgment directly to model development.

Bottom line: Will AI replace doctors?

AI may reduce the administrative workload carried by physicians, but current workforce data doesn't suggest that it will eliminate the need for doctors.

The Association of American Medical Colleges projects a shortage of up to 86,000 physicians by 2036. These figures point to a future where AI assists doctors in managing demand rather than making their roles obsolete.

Responsibility still belongs to physicians

When an AI-assisted recommendation contributes to a poor clinical outcome, liability still remains with the physician and the healthcare system, not the model. That reality shapes how AI gets used in practice.

A tool that can't be held accountable can't be handed the final decision, no matter how good its benchmark scores look.

Medicine requires adapting to uncertainty

The central question is not when AI will replace physicians. It's how fast AI's growing capabilities can be safely absorbed into a healthcare system that needs more clinical capacity, not fewer physicians.

Benchmark scores show that frontier models already perform well on many clinical tasks, and that makes them the ideal tool for transforming how doctors work, not replacing them.

Benchmarks evaluate defined cases, but physicians routinely encounter undefined ones, whether that's novel drug interactions or atypical symptom presentation. This is where physicians will still earn their role, even as benchmark scores improve.

Evaluate medical AI with real-world benchmarks

If your team is building medical AI systems that require clinical expertise, Mercor can connect you with specialists to help you scale.

For teams that need more:

Explore Mercor