AI in medicine: Use cases, benchmarks, & reality

AI in medicine: Use cases, benchmarks, & reality
  • AI in medicine is already used in fields such as imaging, documentation, and drug discovery. However, its adoption is far more mature in pattern-recognition tasks than in those involving judgment-intensive clinical reasoning.
  • Licensing-exam pass rates are a poor indicator of real clinical competence. Benchmarks based on actual physician tasks tell a more honest, and less flattering, story.
  • The concrete benefits of AI, such as faster interpretation, reduced documentation time, and earlier detection, are real, but they come from narrowly focused tools rather than general-purpose assistants.
  • Hallucination, automation bias, and training-data bias aren't edge cases. They're inherent challenges that any clinical AI deployment must be designed to address.
  • For most health systems, the best near-term approach isn't a single AI rollout. It's a series of small, measured pilots that only become part of the standard workflow once they've proved their effectiveness.

What is AI in medicine?

AI in medicine is the use of machine learning, deep learning, natural language processing, and generative AI models to support diagnosis, treatment planning, documentation, and biomedical research. Example applications include a radiology model capable of flagging a suspicious nodule to a drafting tool that turns patient visit notes into clinical documentation.

The adoption of AI in medicine is no longer a fringe issue. The FDA's list of AI-enabled medical devices now includes recent approvals such as Siemens Healthineers' Syngo Carbon and Anumana's ECG-AI Pulmonary Hypertension 12-Lead algorithm, both authorized through the agency's premarket review process.

This list is updated periodically but isn't fully comprehensive, which is itself telling. Regulators can't keep up with the rapid pace at which AI-enabled devices are entering medical practice.

How AI is used in medicine

AI in healthcare falls into three broad categories: interpreting images and signals, supporting clinical decisions and paperwork, and accelerating research into drugs or devices. Each of these areas contributes differently to clinicians' workflows and patient outcomes, and each has a different tolerance for error.

Diagnostics and medical imaging

This is the area where clinical AI is most mature because tasks mainly involve pattern recognition on a fixed input instead of open-ended reasoning. Example use cases include:

  • Radiology: AI models can analyze X-rays, CT scans, and MRI images and flag fractures and nodules for radiologists to review. They're also increasingly capable of identifying hemorrhages that require rapid attention.
  • Pathology: AI can scan digitized tissue slides for cancer markers and cell abnormalities.
  • Dermatology: AI aids in triaging suspicious skin lesions from photographs.
  • Ophthalmology: AI can screen retinal images for diabetic retinopathy and related diseases.

For patients, these uses of AI in medicine usually mean faster triage and earlier detection. For the clinician, it acts as a second pair of eyes that never gets tired, though the final diagnosis still sits with the physician.

Clinical decision support and documentation

In this area, AI serves two distinct functions. One reduces administrative load. The other informs clinical judgment, which carries greater consequences if things go wrong. Common applications include:

  • Triage algorithms: These sort incoming patients or messages by urgency.
  • Risk prediction: AI flags patients at risk of deterioration, readmission, or developing sepsis.
  • Ambient documentation: AI listens to patient visits and drafts clinical notes.
  • Health record summarization: AI condenses lengthy medical records into concise pre-visit summaries.

Documentation tools are one of the most established successes in this category. However, clinical decision-support tools that suggest a diagnosis or next step demand more scrutiny because while a wrong note is merely an annoyance, a wrong recommendation can impact patient safety.

Drug discovery and research

Typically, research-based AI advances faster than patient applications because the cost of a wrong prediction is often just a wasted experiment or compute cycle and doesn't impact patient outcomes. Current uses of AI in research span:

  • Target identification: AI can identify proteins or pathways associated with disease.
  • Molecule screening: It can predict which candidate compounds are worth synthesizing.
  • Clinical-trial optimization: AI can help match eligible patients with clinical trials and predict trial dropout rates.

Because AI errors get caught in a laboratory before they can impact real people, developers can iterate faster and take more risk than those building diagnostic tools.

How Mercor trains AI for medicine

Every AI tool mentioned above depends on expert human judgment somewhere upstream, usually in the training and evaluation phase.

Mercor partners with credentialed physicians to build and grade the clinical data that teaches AI models what a correct diagnosis, safe recommendation, or well-written note actually looks like. This is how models learn from expert feedback and is what separates a model that merely sounds convincing from one that actually performs reliably.

Without physicians in the loop, a model trained primarily on general internet text lacks the context to know the difference.

What are the benefits of AI in medicine?

AI in medicine can help reduce administrative burden, catch diseases earlier, and expand operational capacity. Practical gains include:

  • Time savings on documentation: Ambient scribing tools reduce the time clinicians spend typing notes after a visit, allowing them to dedicate more time to patients rather than paperwork.
  • Earlier detection in screening populations: AI models built for specific diseases, such as diabetic retinopathy, can catch cases that a busy clinic might otherwise miss on a first pass.
  • Proven usability in complex clinical contexts: A randomized, multicenter study in the U.K., which evaluated AI-driven decision support for uncertain antimicrobial prescribing across 23 hospitals and 42 clinicians, reported a system usability scale score of 72.3 out of 100. This demonstrates that a well-designed tool can perform effectively outside of a laboratory setting.
  • Operational throughput: Triage and risk-prediction tools help stretched hospitals prioritize sicker patients, which matters more in systems running near capacity than in well-resourced clinics.

However, none of these benefits are universal. They apply to the specific workflow a tool was designed and validated for, and their effectiveness tends to diminish if stretched beyond their intended use.

What Mercor’s APEX benchmarks reveal about AI’s real clinical reasoning

When implementing AI in medicine, healthcare leaders need to know whether an AI system can support real clinical work, not simply answer medical questions correctly.

The most useful evaluation will ultimately reflect an organization’s own patient population, clinical workflows, and standards of care. Until teams conduct that tailored evaluation, Mercor’s APEX results provide a useful indication of how today’s models perform under realistic clinical demands.

Mercor’s APEX benchmarks evaluate models on 100 realistic primary care tasks developed and scored by credentialed physicians. Models must interpret clinical information, weigh competing factors, and produce recommendations, such as an appropriate medication adjustment. Across these tasks, model performance remains inconsistent.

Scores collected in February 2026 ranged from 46.7% to 71.6%, with only nine out of 19 models scoring at least 60%. Opus 4.6 led the group with a score of 71.6%.

The benchmarks show that frontier models can complete a significant portion of clinical reasoning tasks but still make too many errors to operate independently across a workflow.

The takeaway for healthcare organizations isn't that AI lacks value. Instead, current models may assist with defined clinical tasks, but their reliability depends heavily on what they're being asked to do.

Learn how Mercor can help you benchmark AI in medicine against your organization’s clinical workflows before you commit to a model or platform.

How to implement AI in healthcare

Adopting AI in a clinical setting works best when you follow a series of steps. Skipping these steps is where most deployments tend to fail.

Step 1: Start with high-impact clinical workflows

Begin with high-volume, well-defined tasks, not the most complex problem in your department. Documentation, imaging triage, and administrative routing have clear inputs and outputs, which makes them easier to validate.

Judgment-intensive care, the kind that involves weighing a patient's full history and context, still needs a clinician in the loop, and that's unlikely to change anytime soon.

Step 2: Build on high-quality clinical data

A model is only as reliable as the data it was trained and evaluated on. Well-governed clinical data includes representative patient populations, accurate labels checked by qualified reviewers, and documented provenance.

Cut corners here and a model's apparent accuracy in testing is unlikely to hold up when deployed on a real patient population that looks different from the training data.

Step 3: Choose AI tools clinicians can trust

Match the tool to the specific task it was validated for, not the task you wish it did. For example, a model approved for diabetic retinopathy screening may not be validated for general diagnostic reasoning.

Clinician oversight remains essential because the tool's confidence level doesn't always equate to its correctness, and only a trained professional can identify that gap.

Step 4: Protect patient privacy and compliance

Before adopting any system, ask where patient data goes, who can access it, and whether the vendor's claims about de-identification stand up under scrutiny. These aren't just administrative concerns. They determine whether the deployment is legal, and whether patients would consent to it if asked.

Step 5: Pilot, measure, and expand carefully

Start with a limited workflow, measure results against a real baseline rather than an assumed one, and only scale after the tool demonstrably outperforms that baseline.

The APEX benchmark family is one way to evaluate a model's capability across professional domains before committing budget and clinical time to a broader rollout.

The risks and limits of AI for doctors

The risks associated with clinical AI aren't hypothetical. They're well documented in research, and they shape what a clinician should and shouldn't trust a tool to do. Risks include:

  • Hallucination and factual errors: AI models can generate confident, wrong answers. A Scientific Reports study documented this behavior even in newer, high-performing models.
  • Automation bias: Under time pressure, clinicians can start over-relying on a tool's output, reducing the independent verification that made the tool useful in the first place.
  • Training-data bias: Models trained on unrepresentative patient data perform worse for the populations they saw least during training.
  • Patient privacy: Any tool processing clinical data creates an opportunity for data breaches or misuse.
  • Unclear liability: When an AI system contributes to a poor outcome, who's accountable (whether the clinician, the health system, or the vendor) is often debatable.

How is AI regulated in medicine?

AI in medicine is regulated as software, and that regulatory oversight is getting stricter rather than more lenient.

In the U.S., the FDA issued draft guidance in January 2025, for AI-enabled device software functions. The guidance emphasizes lifecycle management and marketing-submission recommendations rather than treating approval as a one-time event.

The agency has also requested public comment on how to measure the real-world performance of AI-enabled devices after they're deployed, indicating that a good benchmark score at launch isn't considered sufficient on its own.

The FDA also reports publishing 73 new examples of marketing authorizations using real-world evidence between FY2020 and FY2025, including examples supporting AI-enabled technologies. This trend underscores the growing importance of postmarket data in regulatory decision-making.

In the EU, the AI Act classifies most medical AI systems as high-risk, which requires mandatory conformity assessments and ongoing monitoring obligations. Neither regulatory framework considers a strong benchmark score as the final goal.

What comes next for artificial intelligence in medicine?

The near-term trajectory of AI in medicine will be shaped by what's already being tested in clinical practice, not by speculation. Here are some of the key trends emerging in clinical AI usage:

  • Agentic clinical workflows: AI systems will complete multistep tasks, but this will only happen once reliability on real tasks improves well beyond today's ~65% benchmark ceiling.
  • Multimodal models: These models will work across images, notes, and lab results combined, but this is contingent on validation across diverse patient data.
  • Ambient documentation at scale: This area will expand as modest yet real efficiency gains are demonstrated across different specialties.
  • Prospective clinical trials of AI systems: Testing AI tools on live patients rather than retrospective data is becoming the standard of evidence that regulators increasingly expect.

Will AI replace doctors?

AI is replacing specific clinical tasks but not taking over the physician's role, and the current benchmark evidence helps to explain why.

Today's models perform well on narrow, pattern-recognition tasks. A model can read a scan, draft a note, or flag a risk score competently. However, physicians are required for judgment-intensive work, such as interpreting an ambiguous patient history, managing uncertainty, or taking responsibility for decisions.

That's why AI models are unlikely to replace human professionals any time soon.

This same clinical judgment is also what trains and evaluates AI models in the first place. New roles emerging around AI, along with platforms such as Mercor, give physicians and other specialists new opportunities to contribute their expertise directly.

Turn medical AI benchmarks into action

If you're evaluating a new AI tool or planning a clinical rollout, objective benchmark evidence can be invaluable.

For teams that need more:

Learn more about Mercor