Why AI is good at diagnosis but bad at treatment decisions
Four pieces here converge on one structural claim: AI's strength on diagnosis doesn't transfer to treatment, and it isn't a matter of waiting for a better model. Diagnosis rewards exactly what current AI does well — matching new cases against large sets of prior ones. Treatment decisions ask for something else, a judgment call across a patient's competing priorities that doesn't have a single correct answer to pattern-match toward. Taken together, the collected pieces move from explaining that distinction, to the harder ethical question of whether AI should make treatment calls at all, to direct accuracy comparisons between the two tasks, and finally to a study showing how the clinician's own expertise changes what AI assistance adds.
pattern-matching versus values-laden judgment · why diagnosis is a narrower task than treatment · the recurring "AI matches doctors" headline cycle · multi-variable decision-making in treatment planning · why better models don't close this
The study, published in Nature Medicine as 'Divergent impacts of explainable AI for dermatological diagnosis on…
Key takeaways· 2▼
The accuracy boost non-experts got from AI explanations came from deference rather than better reasoning — when the model was wrong, they followed it, and the fluent written explanations inflated confidence most.
The design choice that changed the finding was feeding deliberately wrong AI answers to both groups; a study that only measures performance when the AI is right will overstate the tool's value.
When AI diagnostic advice was incorrect, physician diagnostic accuracy collapsed to 23.6%–26.1% (vs. 85.3%–92.8% when…
Key takeaways· 3▼
Physicians in a Johns Hopkins-affiliated study diagnosed correctly 92.8% of the time when the AI advice was right and 23.6% of the time when it was wrong — a 69-point swing set by the software, not the doctor.
Medical AI names the disease far better than it picks the treatment: one benchmarked system scored 89.3% on diagnosis and 53.0% on treatment, and that gap runs roughly 13 to 47 points across independent studies.
In one pathology study, doctors abandoned a correct diagnosis after a wrong AI suggestion in 7 of every 100 cases — errors that in pathology can mean cancer staged wrong.
The evidence currently supports AI as a powerful adjunct that measurably improves human decision-making in bounded…
Key takeaways· 3▼
The real split over AI-chosen treatments is not accuracy: it is whether picking a treatment is a bigger-data problem AI will solve, or a values judgment where more data never answers what a patient should want.
Doctors who reject a correct AI treatment recommendation face equal or greater malpractice exposure than doctors who follow an incorrect one, which shifts real decision authority toward AI with no rule changing.
The 2018 failure of the cancer-treatment system Watson for Oncology traced not to bad statistics but to a small number of physicians' preferences encoded as if they were validated outcome data.
Diagnosis is a closed-set classification problem with a checkable answer; treatment is an open-ended optimization over…
Key takeaways· 3▼
An AI that reads a retina scan can be graded against biopsy pathology; an AI that recommends a chemotherapy regimen has no reference standard to grade it against — the validation simply does not transfer.
Treatment-recommendation AI is usually trained on the case histories of a limited set of clinicians at one hospital, then reproduces their choices as if they were an objective standard of care.
Whether an AI treatment recommendation was appropriate may only become clear months or years later, entangled with adherence and later treatment changes, so errors take far longer to detect than in diagnosis.
Widely repeated: consumer AI health tools are about as safe as asking a doctor. A Mount Sinai evaluation found OpenAI's health chatbot under-triaged 51.6% of gold-standard emergencies while acing textbook ones.
Key takeaways· 2▼
Doctors are not a reliable backstop for AI errors: among 223 physicians given incorrect AI patient classifications, the wrong labels shaped treatment dosing regardless of how much they said they distrusted AI.
The Science study behind the AI-beats-doctors headlines also reported the model did not significantly outperform physicians on cannot-miss diagnoses, and AI's accuracy edge disappears on treatment suggestion.