The Case For(4)
AI can integrate more data points more consistently than any clinician, and consistency itself is a form of safety in treatment selection
Reasoning: Treatment decisions increasingly depend on synthesizing genomic, imaging, lab, and longitudinal EHR data at a volume no single clinician processes reliably across every patient interaction; AI's advantage is not superior reasoning but superior recall and integration under fatigue and time pressure.
Evidence: Johns Hopkins/Microsoft Azure AI work trained algorithms on electronic health records, imaging, and genomic information to predict disease progression and treatment response, illustrating the data-integration case for AI-driven treatment-relevant prediction.
Moderate strength
In narrow, protocol-governed treatment domains, standalone AI has already demonstrated performance close to or exceeding baseline clinical practice
Reasoning: Where treatment selection reduces to applying a well-validated decision rule (e.g., triage thresholds, dosing algorithms) rather than open-ended judgment, the case for AI-first execution rests on the same logic that supports algorithmic protocols already used in medicine, extended to AI substrates.
Evidence: Nearly 400 FDA-cleared AI algorithms exist for radiology use, most in narrowly defined, well-validated diagnostic tasks — establishing regulatory precedent for AI operating with high autonomy within tightly bounded clinical scopes.
Moderate strength
Liability research suggests physicians who reject correct AI treatment recommendations face equal or greater legal exposure than those who follow them, which functionally already pushes decision authority toward AI
Reasoning: If the incentive structure penalizes clinicians more for overriding accurate AI than for deferring to it, the practical locus of decision-making shifts toward AI even without a formal change in who is declared 'first-line,' and formalizing that shift would only make explicit what already happens implicitly.
Evidence: A 2026 randomized vignette study found that following AI advice did not appear to increase malpractice risk across both jury-based and expert-based legal systems, and other analyses report lay jurors are more likely to hold clinicians liable for rejecting correct AI recommendations than for following incorrect ones.
Contested strength
AI-first treatment decision-making could address access gaps where clinician capacity is the binding constraint
Reasoning: In settings with clinician shortages, a validated AI system making first-line treatment calls with a defined escalation pathway to human review could deliver treatment to more patients faster than a clinician-bottlenecked system, provided the AI's error profile is not worse than the counterfactual of delayed or absent care.
Evidence: Industry commentary frames increasingly autonomous AI systems as a way to address healthcare access challenges in underserved areas, though this remains a projected rather than demonstrated use case for treatment (as opposed to diagnostic) decisions.
Contested strength
The Case Against(6)
Treatment selection is a values-laden trade-off problem, not a prediction problem, and AI systems are architected to predict outcomes, not to weigh a patient's values against risk
Reasoning: Diagnosis asks 'what is true about this patient's state' — a prediction task AI is structurally suited to. Treatment asks 'given uncertainty and this specific patient's risk tolerance, comorbidities, and life context, what should be done' — a normative judgment task that pattern-matching against historical outcomes does not resolve, because the historical cohort's preferences are not the current patient's preferences.
Evidence: The 2018 IBM Watson for Oncology case is instructive precisely because the failure was not raw predictive accuracy but that the system's recommendations were driven by a small number of physicians' treatment preferences encoded into synthetic training cases rather than validated real-patient outcomes — a values-substitution problem, not a data problem.
Strong strength
The liability and accountability architecture for autonomous AI treatment decisions does not exist, and current doctrine actively breaks down when applied to AI-generated treatment plans
Reasoning: Malpractice negligence, hospital vicarious liability, and manufacturer product liability are built around identifiable human decision-makers; when the treatment plan is AI-generated rather than AI-assisted, none of the three doctrines cleanly assigns responsibility for a harmful outcome.
Evidence: Legal scholarship on algorithmic authority describes a resulting "doctrinal collapse" in which the once-clear boundaries between malpractice, vicarious, and product liability disintegrate when diagnostic or treatment reasoning is co-generated by algorithm and physician.
Strong strength
Standalone AI has already underperformed AI-assisted humans on the exact class of task (pattern-recognition-heavy detection) where AI's case is strongest, undercutting the extrapolation to the harder task of treatment selection
Reasoning: If standalone AI cannot yet outperform human-plus-AI teams on bounded, well-defined detection tasks, the argument that AI should be trusted as first-line decision-maker on the more open-ended, higher-stakes task of treatment selection is weaker than proponents claim.
Evidence: In a prospective multicenter diagnostic-accuracy study of 3,409 brain CT scans across 67 Moscow medical organizations, AI-assisted radiologists statistically significantly outperformed standalone AI services on both sensitivity (98.91% vs. 95.91%) and specificity for intracranial hemorrhage detection (p<0.001).
Strong strength
Real-world deployment failures show validated offline performance does not reliably transfer to safe autonomous clinical decisions, and this gap is a known, recurring pattern rather than a one-off
Reasoning: Multiple independently documented cases across different AI systems and clinical domains show the same failure mode — strong development-time metrics collapsing under real-world distribution shift or hidden training-data flaws — suggesting the risk is structural to how these systems are built and validated, not a fixable bug in any one product.
Evidence: Beyond Watson for Oncology, the Epic Sepsis Model — widely deployed in hospitals — showed substantially weaker real-world performance than expected in independent evaluations, illustrating dataset shift as a recurring failure mode distinct from the Watson case.
Strong strength
The current regulatory consensus explicitly rejects autonomous AI treatment decision-making as a category, meaning the proposition would require dismantling rather than extending the existing framework
Reasoning: Regulators have had direct visibility into AI treatment-support tools for years and have consistently drawn the line at clinician-interpreted output rather than autonomous action, which is informative about where the current weight of expert regulatory judgment sits even as the technology has improved.
Evidence: As of the FDA's January 2026 guidance cycle, no autonomous AI prescription services had been cleared by the FDA, and the agency's draft and final guidance both anchor CDS classification on whether the human retains and exercises the actual clinical decision.
Strong strength
Automation bias means clinicians tasked with 'reviewing' AI-first treatment decisions may not meaningfully catch errors, making the safeguard illusory
Reasoning: If the review layer is psychologically compromised by the same deference dynamics that make AI-first decision-making attractive in the first place, the proposed safety architecture (AI decides, human reviews) may not function as intended, especially under time pressure or high caseload.
Evidence: Research on radiologists found that some clinicians accept AI recommendations even when those recommendations are clearly wrong, and errors in AI outputs have been shown to systematically shape physician judgment in ways that persist despite contradicting clinical evidence.
Moderate strength