Brief
The accuracy picture for AI in medicine bifurcates sharply between diagnosis and treatment recommendation, and the split is consistent across multiple independent study designs. On discrete diagnostic tasks, general-purpose large language models now perform at or above published human benchmarks in several settings: GPT-4V reached 88.7% accuracy vs. 77.8% for human readers on NEJM Image Challenge cases, and 73.3% vs. 63.6% on JAMA cases, with GPT-4V also outperforming physicians on the subset of cases physicians answered incorrectly, though the same study flagged that GPT-4V's stated diagnostic rationale was flawed in a notable share of even its correct answers. A separate NEJM AI analysis found GPT-4 correctly diagnosed 57% of cases, outperforming 99.98% of simulated human readers generated from online answers, though the authors themselves cautioned this used a poorly characterized comparator population and likely represents a best-case read for the model. A real-world, non-benchmark comparison at a Cedars-Sinai virtual urgent-care clinic found AI recommendations were rated as optimal more often (77%) than those of physicians (67%) and were less frequently potentially harmful, though the study's authors noted a limitation: they could not measure how much physicians actually relied on the AI output in that setting.
Treatment recommendation accuracy is consistently lower and more brittle than diagnostic accuracy within the same models and same studies — the central asymmetry this brief is built to surface. In a fundus-imaging study cited within a neuroradiology accuracy paper, GPT-4 achieved 55.8% accuracy on a diagnosis task using fundus images and 42.7% on a treatment recommendation task — a roughly 13-point same-model, same-study drop moving from diagnosis to treatment planning. The Doctorina MedBench evaluation of agent-based medical AI found a similar directional pattern but with an added complication: AI Doctor reached a Diagnosis Accuracy of 89.3%, surpassing the 84.6% achieved by GPT-5 and Treatment Accuracy at 53.0%, compared to 38.0% for the baseline — meaning treatment accuracy for the better-performing system (53.0%) was still roughly 36 points below its own diagnostic accuracy (89.3%). Neuroradiology-specific work found comparable degradation: GPT-4 Turbo achieved a baseline diagnostic accuracy of 55.1% on primary imaging diagnosis, in a domain where the same source describes current misdiagnosis rates ranging from 30–50% for LLM-assisted imaging generally — a range wide enough that it should be read as a documented span across studies, not a single point estimate. A dental-medicine exploratory study similarly found AI-generated treatment recommendations aligning with the evidence base at up to 75% agreement, with the best-performing tool (GPT-4) reaching a Cohen's kappa of 0.42 against clinical guidelines — a moderate, not strong, agreement level by standard interrater-reliability thresholds.
The second load-bearing figure set concerns clinician deference to incorrect AI output — automation bias — which is now documented across multiple independent clinical domains with converging but not identical magnitudes. In pathology, a study found in 7 out of every 100 cases where a pathologist had initially reached the correct conclusion, the introduction of an erroneous AI recommendation caused them to abandon that correct answer, and the source explicitly frames this as a patient-safety metric, not statistical noise, given the stakes of missed cancer staging. A related internal-medicine study spanning residents, faculty, and students found the study, encompassing 216 physicians... showed improved diagnostic accuracy, with a more significant increase observed for the students than for senior residents and faculty. However, the study noted that in 6% of cases, clinicians changed their own correct diagnosis in favor of the inaccurate recommendation from the DSS. A wound-maceration simulation with 223 physicians and nurses generating 1,338 decisions found that correct AI assistance proved to be the most influential predictor, with participants having tenfold higher odds of making the correct diagnostic decision when receiving a correct recommendation (OR = 10) — an odds ratio, not a probability, meaning the underlying baseline accuracy rate still governs the absolute risk. In musculoskeletal MRI reading, one laboratory study found that 45.5% of the total mistakes made by clinicians in the AI-assisted round were due to following incorrect AI recommendations, and the same review noted this effect held across all levels of clinician seniority. A 2026 randomized trial registered as NCT06963957 tested whether AI training itself protects against this bias, and found it does not eliminate it: physicians exposed to flawed LLM advice scored worse on overall diagnostic reasoning: 73.3% versus 84.9% in the error-free group, with an adjusted difference of -14.0 percentage points, and top-choice diagnosis accuracy also dropped, from 90.5% in the control group to 76.1% in the flawed-advice group, an adjusted difference of -18.3 percentage points. The consistent finding across pathology, radiology, internal medicine, and now AI-literacy-trained physicians is that formal training reduces but does not eliminate deference to wrong AI output — a structural finding rather than a single-study artifact, replicated across at least five independent research groups and clinical domains cited here.
The regulatory backdrop for interpreting these accuracy figures is itself quantifiable and dated. As of early 2026, the FDA had authorized over 1,350 AI-enabled devices (about double the number in 2022), with the FDA database [showing] a total of 1451 devices (up from 1250 last year)... As of March, 2026, no device has been authorized that uses generative AI or is powered by large language models. This is a critical scope caveat for the entire benchmark literature above: nearly all of the GPT-4/GPT-5/LLM diagnostic and treatment-accuracy figures cited in this brief describe research or investigational use, not any FDA-cleared clinical product — the cleared device base remains dominated by narrower, task-specific machine learning (imaging triage, detection, measurement), concentrated in radiology, which accounted for 74.4% of 2024's newly authorized Class II ML devices. Separately, transparency in the cleared-device evidence base is thin: among 2024 authorizations, only 49 devices (29.2%) reported both sensitivity and specificity; 15.5% provided demographic data — meaning that for roughly seven in ten newly cleared AI devices in 2024, the FDA summary itself does not disclose the sensitivity/specificity figures a clinician would need to calibrate trust in the tool's diagnostic output.
Two structural caveats bound all of the above. First, nearly every diagnostic-accuracy figure here comes from retrospective benchmark testing against curated case sets (NEJM/JAMA quiz archives, radiology report databases, simulated vignettes) rather than prospective, real-world clinical deployment with patient outcomes — these are test-set accuracy rates, not validated clinical-outcome improvements, and Phase II-style benchmark performance historically overstates real-world performance. Second, the automation-bias figures (6%, 7%, 45.5%, the 14–18 percentage-point degradations) come from distinct study designs, patient populations, and AI systems (CNNs, DSSs, and LLMs are not interchangeable failure modes) — they should be read as a consistent directional finding across the literature, not as a single unified error rate.
The Numbers (14)
Physician diagnostic accuracy when AI advice was correct vs. incorrect (same clinicians, same study)
92.8% (correct AI, local explanation) vs. 23.6% (incorrect AI, local explanation)
This is the single clearest documented demonstration that AI output correctness, not clinician skill, drives outcome variance — a ~69 percentage-point swing tied to whether the AI happened to be right.
As of reported November 2024ScienceDaily coverage of Johns Hopkins-affiliated studyMedium confidence
GPT-4V diagnostic accuracy vs. human readers, NEJM Image Challenge
88.7% (GPT-4V) vs. 77.8% (human readers)▲ Up
An 11-point advantage for the AI model on a curated quiz benchmark; NEJM Image Challenge cases are a controlled test set, not real-world diagnostic complexity.
As of study posted November 2023 (medRxiv)Comparative Analysis of GPT-4Vision, GPT-4 and Open Source LLMs in Clinical Diagnostic Accuracy (medRxiv preprint)Medium confidence
GPT-4V diagnostic accuracy vs. human readers, JAMA Clinical Challenge
73.3% (GPT-4V) vs. 63.6% (human readers)▲ Up
Consistent ~10-point AI advantage on a second independent quiz benchmark, reinforcing the NEJM finding rather than standing alone.
As of study posted November 2023 (medRxiv)Comparative Analysis of GPT-4Vision, GPT-4 and Open Source LLMs in Clinical Diagnostic Accuracy (medRxiv preprint)Medium confidence
GPT-4 correct diagnosis rate on complex NEJM case challenges
57%, outperforming 99.98% of simulated human readers
The authors themselves caveat this as a likely best-case estimate because the human comparator was a poorly characterized population of online quiz-takers, not credentialed physicians under clinical conditions.
As of published in NEJM AI, cited via ai.nejm.orgNEJM AI (Kanjee, Crowe, Rodman; also published in JAMA 2023;330:78-80)Medium confidence
AI diagnosis accuracy vs. treatment-recommendation accuracy, same model (fundus imaging)
55.8% (diagnosis) vs. 42.7% (treatment recommendation)▼ Down
A ~13-point same-model drop moving from diagnosis to treatment planning — direct evidence for the diagnosis-vs-treatment accuracy gap this scan is designed to measure.
As of cited within a 2024 neuroradiology accuracy paper (PMC11276551)GPT-4 fundus imaging study, cited secondarily in PMC11276551Low confidence
Diagnosis Accuracy vs. Treatment Accuracy, AI Doctor system (Doctorina MedBench)
89.3% diagnosis vs. 53.0% treatment (AI Doctor); 84.6% diagnosis vs. 38.0% treatment (GPT-5 baseline)
For both systems tested, treatment accuracy trails diagnostic accuracy by roughly 36-47 percentage points — the largest documented diagnosis-to-treatment gap in this dataset, though from a single unreplicated benchmark.
As of arXiv preprint dated 2026Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI (arXiv)Low confidence
LLM-assisted neuroradiology misdiagnosis rate range
30%–50%
A documented span across studies rather than a single point estimate; the same paper's own baseline GPT-4 Turbo accuracy (55.1%) sits within this error-rate range, underscoring imaging-specific difficulty.
As of cited in 2024 study (PMC11276551)Optimizing GPT-4 Turbo Diagnostic Accuracy in Neuroradiology (PMC)Medium confidence
Pathologists abandoning a correct diagnosis after erroneous AI suggestion
7 per 100 cases (7%)
Framed explicitly by the source as a patient-safety figure given that pathology errors can mean missed cancer staging, not a statistically negligible rounding effect.
As of cited in The Pathologist, April 2026 coverageRosbach study, cited in The PathologistMedium confidence
Physicians switching correct diagnosis to incorrect DSS recommendation
6% of cases
A second, independent domain (general internal medicine, spanning students through faculty) replicating a similar-magnitude override rate to the pathology figure above.
As of cited in 2024 Bowtie-analysis review (ScienceDirect)Friedman et al., 216-physician study, cited in ScienceDirect automation-bias reviewMedium confidence
Clinician mistakes attributable to following incorrect AI advice, MRI/ACL reading
45.5% of total AI-assisted-round mistakes
The highest documented automation-bias attribution rate among the studies compiled here, and the source notes this held across all clinician experience levels.
As of cited in ICE Blog, August 2025Laboratory study on AI-assisted ACL rupture MRI diagnosis, cited in ICE BlogLow confidence
Diagnostic reasoning accuracy, AI-trained physicians given flawed vs. error-free LLM advice
73.3% (flawed advice) vs. 84.9% (error-free), adjusted difference -14.0 percentage points▼ Down
Demonstrates that formal AI training does not eliminate automation bias — physicians with dedicated AI training still showed a clinically meaningful, statistically adjusted drop in reasoning quality when advice was wrong.
As of randomized trial NCT06963957, reported May 2026BJH coverage of NCT06963957 randomized trialMedium confidence
Top-choice diagnosis accuracy, same trial, flawed vs. error-free LLM advice
76.1% (flawed advice) vs. 90.5% (error-free), adjusted difference -18.3 percentage points▼ Down
The largest single-study percentage-point degradation in this dataset, on a trial specifically designed to test whether AI literacy training is protective — the answer is only partially.
As of randomized trial NCT06963957, reported May 2026BJH coverage of NCT06963957 randomized trialMedium confidence
FDA-authorized AI/ML-enabled medical devices, cumulative
~1,350–1,451 devices▲ Up
Roughly double the 2022 count; as of March 2026 no device on this list uses generative AI or LLM technology, meaning the FDA-cleared device base and the LLM benchmark literature above describe two largely non-overlapping populations of AI systems.
As of early 2026 / March 2026IntuitionLabs analysis and Medical Futurist, both citing the FDA AI/ML device listMedium confidence
2024-authorized ML devices reporting both sensitivity and specificity
29.2% (49 of 168 devices)▬ Flat
Roughly seven in ten newly cleared devices in 2024 did not disclose sensitivity/specificity in FDA summaries, limiting clinicians' ability to independently calibrate trust against the accuracy benchmarks reported elsewhere in this brief.
As of 2024, published in a 2026 cross-sectional analysis (PMC12730494)Machine Learning-Enabled Medical Devices Authorized by the US FDA in 2024 (PMC)Medium confidence
Comparisons (4)
AI diagnostic accuracy vs. AI treatment-recommendation accuracy (same model, fundus imaging task)
55.8% diagnosis (GPT-4)vs42.7% treatment recommendation (GPT-4)
Gap: ~13 percentage points lower for treatment planning, the core asymmetry this scan was built to quantify
AI Doctor system: diagnosis vs. treatment accuracy (Doctorina MedBench)
89.3% diagnosis accuracyvs53.0% treatment accuracy
Gap: ~36 percentage points lower for treatment recommendation, the largest gap documented in this dataset
Physician diagnostic accuracy when AI advice is correct vs. incorrect
92.8% (correct AI advice, local explanation)vs23.6% (incorrect AI advice, local explanation)
Gap: ~69 percentage-point swing tied entirely to AI correctness, not clinician skill
AI-trained physicians' top-choice diagnostic accuracy: error-free vs. flawed LLM advice
90.5% (error-free advice)vs76.1% (flawed advice)
Gap: 18.3 percentage-point adjusted drop, demonstrating AI-literacy training reduces but does not eliminate automation bias
Facts & Figures (14)
The claims behind this analysis, each with its verification status — including what is contested, unverified, or could not be established.
Physician diagnostic accuracy when AI advice was correct vs. incorrect (same clinicians, same study): 92.8% (correct AI, local explanation) vs. 23.6% (incorrect AI, local explanation)
This is the single clearest documented demonstration that AI output correctness, not clinician skill, drives outcome variance — a ~69 percentage-point swing tied to whether the AI happened to be right.
— FROM THE RECORDper ScienceDaily coverage of Johns Hopkins-affiliated study · as of reported November 2024 · Medium confidence
GPT-4V diagnostic accuracy vs. human readers, NEJM Image Challenge: 88.7% (GPT-4V) vs. 77.8% (human readers)
An 11-point advantage for the AI model on a curated quiz benchmark; NEJM Image Challenge cases are a controlled test set, not real-world diagnostic complexity.
— FROM THE RECORDper Comparative Analysis of GPT-4Vision, GPT-4 and Open Source LLMs in Clinical Diagnostic Accuracy (medRxiv preprint) · as of study posted November 2023 (medRxiv) · Medium confidence
GPT-4V diagnostic accuracy vs. human readers, JAMA Clinical Challenge: 73.3% (GPT-4V) vs. 63.6% (human readers)
Consistent ~10-point AI advantage on a second independent quiz benchmark, reinforcing the NEJM finding rather than standing alone.
— FROM THE RECORDper Comparative Analysis of GPT-4Vision, GPT-4 and Open Source LLMs in Clinical Diagnostic Accuracy (medRxiv preprint) · as of study posted November 2023 (medRxiv) · Medium confidence
GPT-4 correct diagnosis rate on complex NEJM case challenges: 57%, outperforming 99.98% of simulated human readers
The authors themselves caveat this as a likely best-case estimate because the human comparator was a poorly characterized population of online quiz-takers, not credentialed physicians under clinical conditions.
— FROM THE RECORDper NEJM AI (Kanjee, Crowe, Rodman; also published in JAMA 2023;330:78-80) · as of published in NEJM AI, cited via ai.nejm.org · Medium confidence
AI diagnosis accuracy vs. treatment-recommendation accuracy, same model (fundus imaging): 55.8% (diagnosis) vs. 42.7% (treatment recommendation)
A ~13-point same-model drop moving from diagnosis to treatment planning — direct evidence for the diagnosis-vs-treatment accuracy gap this scan is designed to measure.
— FROM THE RECORDper GPT-4 fundus imaging study, cited secondarily in PMC11276551 · as of cited within a 2024 neuroradiology accuracy paper (PMC11276551) · Low confidence
Diagnosis Accuracy vs. Treatment Accuracy, AI Doctor system (Doctorina MedBench): 89.3% diagnosis vs. 53.0% treatment (AI Doctor); 84.6% diagnosis vs. 38.0% treatment (GPT-5 baseline)
For both systems tested, treatment accuracy trails diagnostic accuracy by roughly 36-47 percentage points — the largest documented diagnosis-to-treatment gap in this dataset, though from a single unreplicated benchmark.
— FROM THE RECORDper Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI (arXiv) · as of arXiv preprint dated 2026 · Low confidence
LLM-assisted neuroradiology misdiagnosis rate range: 30%–50%
A documented span across studies rather than a single point estimate; the same paper's own baseline GPT-4 Turbo accuracy (55.1%) sits within this error-rate range, underscoring imaging-specific difficulty.
— FROM THE RECORDper Optimizing GPT-4 Turbo Diagnostic Accuracy in Neuroradiology (PMC) · as of cited in 2024 study (PMC11276551) · Medium confidence
Pathologists abandoning a correct diagnosis after erroneous AI suggestion: 7 per 100 cases (7%)
Framed explicitly by the source as a patient-safety figure given that pathology errors can mean missed cancer staging, not a statistically negligible rounding effect.
— FROM THE RECORDper Rosbach study, cited in The Pathologist · as of cited in The Pathologist, April 2026 coverage · Medium confidence
Physicians switching correct diagnosis to incorrect DSS recommendation: 6% of cases
A second, independent domain (general internal medicine, spanning students through faculty) replicating a similar-magnitude override rate to the pathology figure above.
— FROM THE RECORDper Friedman et al., 216-physician study, cited in ScienceDirect automation-bias review · as of cited in 2024 Bowtie-analysis review (ScienceDirect) · Medium confidence
Clinician mistakes attributable to following incorrect AI advice, MRI/ACL reading: 45.5% of total AI-assisted-round mistakes
The highest documented automation-bias attribution rate among the studies compiled here, and the source notes this held across all clinician experience levels.
— FROM THE RECORDper Laboratory study on AI-assisted ACL rupture MRI diagnosis, cited in ICE Blog · as of cited in ICE Blog, August 2025 · Low confidence
Diagnostic reasoning accuracy, AI-trained physicians given flawed vs. error-free LLM advice: 73.3% (flawed advice) vs. 84.9% (error-free), adjusted difference -14.0 percentage points
Demonstrates that formal AI training does not eliminate automation bias — physicians with dedicated AI training still showed a clinically meaningful, statistically adjusted drop in reasoning quality when advice was wrong.
— FROM THE RECORDper BJH coverage of NCT06963957 randomized trial · as of randomized trial NCT06963957, reported May 2026 · Medium confidence
Top-choice diagnosis accuracy, same trial, flawed vs. error-free LLM advice: 76.1% (flawed advice) vs. 90.5% (error-free), adjusted difference -18.3 percentage points
The largest single-study percentage-point degradation in this dataset, on a trial specifically designed to test whether AI literacy training is protective — the answer is only partially.
— FROM THE RECORDper BJH coverage of NCT06963957 randomized trial · as of randomized trial NCT06963957, reported May 2026 · Medium confidence
FDA-authorized AI/ML-enabled medical devices, cumulative: ~1,350–1,451 devices
Roughly double the 2022 count; as of March 2026 no device on this list uses generative AI or LLM technology, meaning the FDA-cleared device base and the LLM benchmark literature above describe two largely non-overlapping populations of AI systems.
— FROM THE RECORDper IntuitionLabs analysis and Medical Futurist, both citing the FDA AI/ML device list · as of early 2026 / March 2026 · Medium confidence
2024-authorized ML devices reporting both sensitivity and specificity: 29.2% (49 of 168 devices)
Roughly seven in ten newly cleared devices in 2024 did not disclose sensitivity/specificity in FDA summaries, limiting clinicians' ability to independently calibrate trust against the accuracy benchmarks reported elsewhere in this brief.
— FROM THE RECORDper Machine Learning-Enabled Medical Devices Authorized by the US FDA in 2024 (PMC) · as of 2024, published in a 2026 cross-sectional analysis (PMC12730494) · Medium confidence