The Headline
AI reasoning models now match or beat physicians on structured diagnostic pattern-matching, but the diagnosis-beats-doctors finding does not extend to treatment management, uncertainty navigation, or unsupervised consumer use — and physicians who receive incorrect AI classifications frequently fail to override them even against contradictory evidence.
Overview
Reasoning-model AI has genuinely matched or exceeded physician baselines on curated diagnostic tasks, but the same evidence base shows AI's edge collapses on treatment management, uncertainty navigation, and messy real-world triage — and physicians themselves are poor at overriding AI when it is wrong.
Brief
The claim that "AI beats doctors" traces to a real and methodologically serious finding. In the Science paper by Brodeur et al., OpenAI's o1-preview reasoning model was tested across six experiments against physician baselines, including 76 actual emergency department cases with blinded physician adjudication. On published clinicopathological conference cases, GPT-4 achieved exact or very close diagnostic accuracy in 72.9% of cases, whereas o1-preview achieved this in 88.6%, and in actual emergency department cases, o1 achieved 67.1% exact or very-close diagnostic accuracy at initial triage, outperforming two expert attending physicians at 55.3% and 50.0%. A companion Science perspective piece states plainly that this is an important limitation of existing benchmarks of medical AI... the question the paper answers — can a frontier model match physicians on curated cases — has largely been settled in the affirmative across multiple studies, and that the field's open questions have shifted to the joint human-AI system, the failure modes of that system, and the curation step that vignette-based evaluation cannot capture. That last sentence is the whole story compressed. The diagnostic-matching finding is real, replicated across benchmark types, and methodologically defensible — but it is narrower than the headlines that followed it. The same Science commentary notes two important qualifications inside the flagship study itself: the evaluation relies on curated, text-only inputs, and as the authors acknowledge, existing benchmarks rely on the careful work of clinicians to curate and clean up cases, which may overstate AI performance when using messy data available in more realistic clinical workflows — and even within these controlled settings, the model did not outperform physicians on cannot-miss diagnoses, and its advantage in landmark diagnostic cases was not statistically significant. A separate commentary flags a second confound: the comparator itself has moved, since the AMA's 2026 Augmented Intelligence Research survey found 81% of US physicians now report using AI professionally, up from 38% in 2023 and 66% in 2024 — meaning "AI vs. doctors" increasingly compares AI against doctors who are themselves already using AI. The overreach begins the moment "diagnosis" gets substituted for "management" or "treatment decisions." An earlier systematic review comparing AI and clinician diagnostic performance found that AI transcended the average levels of clinicians in most clinical situations except for treatment suggestion — the exception was already visible before the current reasoning-model generation. Consumer deployment makes the gap concrete and dangerous rather than abstract: an independent Mount Sinai evaluation of OpenAI's consumer-facing ChatGPT Health, published in Nature Medicine, found the tool followed an inverted U-shaped pattern, with the most dangerous failures concentrated at clinical extremes — nonurgent presentations (35%) and emergency conditions (48%), and separately reported that it under-triaged 51.6% of emergency cases, recommending the patient see a doctor instead of going to the emergency room. Critically, emergencies like stroke, with unmistakable symptoms, were correctly triaged 100% of the time — the model's failures cluster precisely where clinical judgment under ambiguity matters most, which is the opposite of where the curated-vignette benchmarks test it. A methodological critique of that same study argues its exam-style protocol — forced A/B/C/D output, knowledge suppression, and suppression of clarifying questions — differs fundamentally from how consumers use health chatbots, so even this negative finding is contested on design grounds rather than settled. The most consequential misconception may not be about AI's capability at all, but about physicians' capacity to safely supervise it. A PLOS Digital Health study of 223 self-reported physicians found that incorrect AI patient classification influenced the majority of participants' decisions independently of their attitudes and perceptions about AI reliability, and most physicians administered more treatment doses to patients incorrectly classified as highly sensitive, with their effectiveness judgments also higher for these patients — even when outcome data directly contradicted the AI. In the more severe of the two experiments, the drug was completely ineffective, yet participants generally judged it to be effective. That finding sits in tension with a separate ICU simulation finding that clinicians are not uniformly credulous: in a physical simulation study, 92% of clinicians rejected unsafe AI recommendations versus only 29% of safe ones, suggesting override behavior depends heavily on how starkly wrong the AI recommendation is, not just whether it is wrong.
Myths & Realities (5)
Myth
AI now beats doctors at diagnosis, full stop — the science is settled.
Reality
The finding is real but narrower than commonly stated: it applies to text-based diagnostic pattern-matching on largely curated cases (plus one blinded ED arm), not to unsupervised, real-world clinical practice with incomplete records and time pressure.
Evidence: The flagship Science study found o1-preview reaching 88.6% accuracy on published case vignettes versus 72.9% for GPT-4, and outperforming attending physicians on ED triage accuracy (67.1% vs 55.3%/50.0%), but a companion perspective notes the model did not significantly outperform physicians on cannot-miss diagnoses and that benchmarks rely on cleaned-up cases that may overstate real-world performance.
Kernel of truth: Reasoning-model AI genuinely does match or exceed physicians on structured diagnostic accuracy metrics across multiple independent studies, including one real-world blinded ED comparison — this is not a fabricated result.
Why believed: Headline coverage of the Brodeur et al. Science study and NPR/Fortune reporting compressed a six-experiment, multi-caveat paper into a single 'AI beat doctors' framing that traveled further than the paper's own qualifications.
Myth
If AI can diagnose accurately, it can also manage treatment and personalize care just as well.
Reality
The evidence base explicitly separates these skills: AI's advantage evaporates on treatment suggestion and management reasoning, which requires navigating uncertainty, patient-specific tradeoffs, and shared decision-making that current models handle poorly.
Evidence: A systematic review found AI diagnostic accuracy exceeded clinician averages in most situations except treatment suggestion; independent reporting on the diagnosis-vs-treatment literature describes physicians as still superior at weighing treatment options and personalized care, a process termed 'management reasoning.'
Kernel of truth: AI models can generate plausible-sounding treatment options and cite relevant literature quickly, which is a genuine efficiency aid even if the final judgment quality lags physicians.
Why believed: Diagnosis and treatment are colloquially treated as one continuous task ('figuring out what's wrong and what to do about it'), so a benchmark win on the first half gets generalized to the whole clinical encounter.
Myth
Consumer AI health tools are essentially as safe as talking to a doctor because they're built on the same accurate diagnostic AI.
Reality
An independent safety evaluation of OpenAI's consumer-facing ChatGPT Health found it under-triaged roughly half of gold-standard emergency cases while performing perfectly on textbook emergencies, showing consumer deployment context and ambiguous presentations — not raw model capability — drive the most dangerous failures.
Evidence: The Mount Sinai Nature Medicine study (960 interactions across 60 clinician-authored vignettes) found 51.6% undertriage of emergencies including diabetic ketoacidosis and impending respiratory failure, alongside 100% correct triage of unambiguous emergencies like stroke and anaphylaxis.
Kernel of truth: The same tool performed flawlessly on unambiguous, textbook-presentation emergencies, so consumer AI health tools are not uniformly unsafe — the failure is concentrated and identifiable rather than random.
Why believed: Marketing and rapid adoption (reported figures cite roughly 40 million daily US health-related queries to the tool) create an impression of validated safety that the underlying research had not yet independently tested at the time of launch.
Myth
Doctors act as a reliable safety net that catches AI's diagnostic mistakes before they reach patients.
Reality
Controlled experiments show physicians frequently fail to override incorrect AI classifications even when directly contradicted by patient outcome data, sometimes concluding an entirely ineffective treatment was working.
Evidence: A PLOS Digital Health study of 223 physicians found incorrect AI patient classifications shaped treatment dosing and effectiveness judgments regardless of physicians' stated trust or skepticism toward AI, and in one experiment physicians did not recognize that a treatment was completely ineffective despite receiving outcome feedback across dozens of decisions.
Kernel of truth: A separate ICU simulation found clinicians rejected unsafe AI recommendations 92% of the time versus only 29% of safe ones, showing override behavior is not absent — it appears to depend on how starkly dangerous the AI's suggestion is.
Why believed: The 'human in the loop' framing is widely promoted as the default safety mechanism for clinical AI, and it is intuitive that trained experts would catch obvious errors, but the empirical override research is recent and less publicized than the capability benchmarks.
Myth
AI diagnostic confidence is a reliable signal of accuracy, similar to physician confidence.
Reality
Confidence in medicine, whether from a clinician or a model, is not a reliable proxy for correctness, and this is a documented human vulnerability that AI-generated confidence can exploit just as effectively as physician overconfidence historically has.
Evidence: Commentary on clinical decision-making cites a finding that ICU physicians who were 'completely certain' of their diagnosis were still wrong a substantial share of the time, illustrating that stated confidence — human or AI-generated — does not track reliably with accuracy in ambiguous clinical situations.
Kernel of truth: Well-calibrated confidence scores can be a useful input when validated against outcomes, and some benchmark studies do compute formal calibration metrics for AI diagnostic confidence.
Why believed: People generally associate fluent, decisive-sounding output with expertise, and AI reasoning models produce especially fluent, well-structured justifications that can read as more authoritative than a hedging physician even when the underlying accuracy is not higher.
The Corrected View
The accurate picture has two separable layers: on structured diagnostic pattern-matching, especially against curated or benchmark-style cases, frontier reasoning models have demonstrated performance matching or exceeding physician baselines in multiple peer-reviewed studies, including one blinded real-world emergency department comparison. On treatment management, personalized risk-benefit weighing, and unsupervised consumer triage of ambiguous presentations, the evidence shows AI still underperforms, sometimes dangerously so in real-world deployment, and that physicians are not automatically an effective backstop against AI's errors. Neither the enthusiastic "AI has surpassed doctors" narrative nor a dismissive "AI can't be trusted for anything clinical" narrative fits the documented record; the responsible reading is task-specific and deployment-context-specific rather than a single verdict on AI-versus-doctors.
Facts & Figures (6)
The claims behind this analysis, each with its verification status — including what is contested, unverified, or could not be established.
OpenAI's o1-preview achieved 88.6% exact/very-close diagnostic accuracy on published clinicopathological conference cases versus 72.9% for GPT-4, and 67.1% accuracy at ED triage versus 55.3% and 50.0% for two attending physicians (Brodeur et al., Science, 2026)
This is the specific, dated, peer-reviewed effect size underlying the 'AI beats doctors at diagnosis' claim — a real finding, not overreach, when confined to diagnostic pattern-matching on this study's design.
✓ GROUNDED
The same Science paper's authors acknowledge existing benchmarks rely on clinician-curated 'cleaned up' cases that may overstate AI performance versus messy real-world clinical workflow data, and the model did not significantly outperform physicians on 'cannot-miss' diagnoses
Directly bounds the diagnostic-matching claim: even the flagship pro-AI study has a stated ceiling and a null result on the highest-stakes diagnostic category.
✓ GROUNDED
An independent Mount Sinai evaluation of ChatGPT Health (Ramaswamy et al., Nature Medicine, 2026) found 51.6% undertriage of gold-standard emergency cases and roughly 65% overtriage of nonurgent cases, with 100% correct triage on unambiguous emergencies like stroke and anaphylaxis
Demonstrates the diagnosis-vs-management/triage gap in a live consumer deployment, showing failure concentrated exactly where ambiguity and clinical judgment matter most.
✓ GROUNDED
A PLOS Digital Health experiment with 223 self-reported physicians found that incorrect AI patient classifications shaped treatment dosing and physicians' effectiveness judgments even when contradicted by patient outcome data, and in one experiment physicians failed to recognize a completely ineffective treatment as ineffective
Shows the risk is not only AI's residual error rate but clinicians' documented difficulty overriding AI even with contradictory evidence in hand — a human-factors failure mode distinct from model accuracy.
✓ GROUNDED
A 2019 systematic review found AI diagnostic accuracy transcended clinician average performance in most clinical situations except treatment suggestion, where clinicians retained an advantage
Shows the diagnosis/treatment performance gap predates the current reasoning-model generation, indicating a structural rather than transient limitation.
✓ GROUNDED
A methodological critique (Navarro et al., 2026) argues the ChatGPT Health undertriage study used an exam-style forced-choice protocol that suppresses clarifying questions, differing from how consumers actually converse with health chatbots
Prevents overclaiming certainty on the negative finding too — even the safety-failure evidence is contested on methodological grounds, illustrating that neither the 'AI is great' nor 'AI is unsafe' framing should be taken as fully settled.
✓ GROUNDED