CRBC News
Health

AI Outperforms Doctors in Multiple Diagnostic Tests — What That Means for Medicine

AI Outperforms Doctors in Multiple Diagnostic Tests — What That Means for Medicine
Illustration: Joanna Andreasson; Source images: iStock

Recent experiments from Beth Israel Deaconess and Harvard Medical School found that OpenAI's preview model o1 outperformed physicians on multiple diagnostic tests—identifying the correct diagnosis in 78% of case studies versus about 30% for doctors, and scoring 89% on clinical vignettes versus 34% for clinicians. Other research—including AI-assisted mammography and Mayo Clinic work on pancreatic cancer—shows earlier or improved detection. Policymakers in states such as Nevada and Illinois are already restricting some clinical uses of AI, which could affect adoption despite promising results.

Artificial intelligence (AI) systems are beginning to outperform physicians on a range of diagnostic tasks, according to multiple recent studies. In experiments led by researchers at Beth Israel Deaconess Medical Center and Harvard Medical School, a preview release of OpenAI's large language model, dubbed o1, exceeded physician performance on several clinical and diagnostic reasoning benchmarks.

Key Study Results

Using clinical case studies, investigators asked o1 to generate lists of possible diagnoses. The model included the correct diagnosis in 78% of cases, compared with roughly 30% for physicians on the same tasks. In a separate assessment using five real clinical vignettes, two physician reviewers scored the model’s recommended next steps: o1 averaged 89% while human clinicians averaged 34%.

The study, published April 30 in Science, also evaluated o1 on emergency department cases where information is often limited. Blinded reviewers found that o1 outperformed an earlier LLM and two attending physicians, particularly during initial triage: o1 produced the exact or a very close diagnosis in 67.1% of cases, versus 55.3% and 50% for the two physicians.

Other Supporting Research

These findings align with additional work showing AI can improve screening and early detection. A Swedish study published in The Lancet reported that AI-assisted mammography improved breast cancer detection. Researchers at the Mayo Clinic trained a model on abdominal CT scans and identified patients who later received pancreatic cancer diagnoses an average of 475 days earlier—and in some cases up to three years earlier—than clinicians recognized, a result reported in the journal Gut.

Context and Cautions

“Humans should be the ultimate baseline,” said Peter Brodeur, a co-author of the Harvard study. AI tools are powerful aids but are not poised to replace clinicians overnight.

Experts emphasize that LLMs and other AI systems can assist diagnostic workflows, flag missed possibilities, and prioritize urgent cases, potentially improving outcomes when integrated with clinician oversight. However, models can still make errors, and rigorous prospective validation, monitoring for bias, and careful integration into clinical workflows are essential.

Regulatory Landscape

Policy responses are already evolving. Nevada enacted a law prohibiting AI systems from implying they can provide professional mental or behavioral health care or from delivering services that would constitute such practice. Illinois passed legislation last year preventing AI from providing therapy and restricting therapists from using AI to make independent therapeutic decisions or to detect emotions or mental states. Other states, including Ohio, California, Minnesota, and Kentucky, are considering related measures—some restrictive, others that would mandate chatbot ability to detect and respond to mental-health needs.

These regulatory choices could slow deployment of AI tools in sensitive clinical areas even as research suggests they can sometimes diagnose conditions more accurately or earlier than clinicians working alone.

Bottom line: Recent evidence shows AI models can match or exceed physician performance on specific diagnostic tests. Careful validation, clinician oversight, and sensible regulation will determine whether that potential translates into safer, more effective care for patients.

Help us improve.

Related Articles

Trending