Medical educators warn that heavy reliance on clinical AI risks trainees "never-skilling"—failing to develop core clinical reasoning. A recent Nature Medicine study found many medical-specific AIs perform worse than general-purpose chatbots like ChatGPT and Claude. Experts recommend periodic AI-free case work, stricter rules for student AI use, and other safeguards to preserve critical thinking.
Doctors Warn Med Students Are Outsourcing Clinical Reasoning To Flawed Medical AIs

Doctors around the world already use artificial intelligence for tasks ranging from note-taking to quick clinical lookups. But clinicians on the front line are raising alarms about how trainees rely on these tools—and how that reliance may be reshaping medical thinking for the worse.
In an essay for The Guardian, Simar Bajaj, a journalist and medical student at Stanford University School of Medicine, and Joseph Sakran, a trauma surgeon and executive vice-chair of surgery at Johns Hopkins Medicine, warn that AI use among trainees risks "not just deskilling but never-skilling." As they put it:
"Although a doctor who has forgotten how to reason is recoverable, one who never learned how may not be."
Two Troubling Findings
The authors point to two recent findings that underscore their concern. Internal data indicate that roughly two-thirds of U.S. doctors are already using OpenEvidence, an AI chatbot designed for clinicians—suggesting the technology is already embedded in clinical workflows.
They also cite a study published in Nature Medicine that evaluated medical-specific large language models. That research found many clinical AI tools answered medical queries less accurately than general-purpose chatbots such as ChatGPT and Claude. According to the study's authors, specialist systems performed roughly on par with Google's AI Overviews, which are known to produce uneven or inaccurate summaries.
Why This Matters
Bajaj and Sakran emphasize that while AI can be a helpful research aid, it cannot replace the cognitive work of learning. "The struggle is the point," they write. When trainees are asked to generate a differential diagnosis themselves, the effort—and the corrective feedback when they miss something—builds clinical reasoning. If a trainee simply queries an AI and receives a near-perfect differential, that formative struggle is lost.
What makes AI different from many past medical advances, the authors argue, is that these systems are not just extending perception or access to data; they are inserting themselves into the "cognitive machinery" training is meant to build.
Evidence and Responses
Research on "cognitive offloading" supports these concerns: studies suggest routinely outsourcing thought to tools can impair critical thinking and reduce brain activity during cognitive tasks. In response, medical educators from top institutions have proposed stricter frameworks for student use of AI.
Practical measures suggested by Bajaj and Sakran include requiring trainees to periodically work through cases without AI assistance—analogous to the Federal Aviation Administration's guidance that pilots occasionally disengage autopilot and fly manually. Social pressures may also play a role: a Johns Hopkins student reported that colleagues who overuse AI were sometimes met with scorn or skepticism.
What To Watch For
As clinical AI tools spread, institutions must balance the efficiencies these tools offer with deliberate training safeguards so that future doctors develop robust clinical reasoning. Without careful policy, assessment, and cultural norms, there is a genuine risk that trainees will outsource thinking they ought to internalize.
Help us improve.

























