Overview: Scott Winters alleges ChatGPT-4o gave reassuring medical advice that preceded a life-threatening pulmonary embolism. Research shows modern AI often excels on structured diagnostic tests, sometimes outperforming clinicians, but struggles in unscripted, real-world interactions. Studies highlight risks of hallucinations and dangerous omissions; one benchmark found direct application risked severe harm in 24.6% of cases. The core issue appears to be deployment and safeguards rather than raw model capability.
Did ChatGPT Delay Critical Care? Lawsuit Against OpenAI Raises Questions About AI Diagnosis

Summary: A lawsuit by Scott Winters alleges ChatGPT-4o gave medical guidance that downplayed serious symptoms and contributed to a life-threatening pulmonary embolism. The case — and a companion suit tied to an overdose death — highlights an urgent question: how reliable are large language models for medical advice outside controlled testing environments?
Case Details
Scott Winters, a former Florida pastor, sued OpenAI and CEO Sam Altman in San Francisco County Superior Court in July 2026. According to the complaint, Winters repeatedly consulted ChatGPT-4o in 2025 after experiencing dizziness and unstable blood pressure. He alleges the chatbot minimized his symptoms, recommended he remain "recliner-bound," and suggested he would need eight to ten similar episodes before his condition warranted real concern.
Weeks later, Winters suffered a massive pulmonary embolism — a blood clot in his lungs. One of his treating physicians reportedly linked the embolism to prolonged immobility. The complaint also cites a specific exchange on the day of the event in which Winters asked whether groin tenderness required an ER visit; the bot allegedly replied,
"God did not design your body to endlessly fail."Hours after that exchange, he nearly died.
OpenAI has responded that ChatGPT is not intended to replace health-care professionals and that its terms of service warn users not to rely on it as their sole source of medical guidance. Winters’ legal team seeks monetary damages and an injunction to pause ChatGPT Health pending an independent safety review.
How This Fits With The Research
Researchers have tested modern large language models (LLMs) on a range of diagnostic and triage tasks. Results are mixed and depend heavily on the testing setup.
Strong Results In Structured Tests
- A 2024 study in JAMA Internal Medicine compared GPT-4 with 21 attending physicians and 18 residents across 20 clinical cases using the validated r-IDEA clinical-reasoning scale; GPT-4 posted a median score of 10/10 compared with 9 for attendings and 8 for residents, although the study noted GPT-4 produced some "flatly incorrect" answers.
- A follow-up in JAMA Network Open reported ChatGPT achieved 90% diagnostic accuracy on six particularly difficult cases when operating alone, while unassisted physicians scored 74% and physicians using ChatGPT as an assistant scored 76%.
Larger-Scale and Meta-Analytic Findings
- A 2025 Nature study testing Google’s AMIE model on 302 complex, real-world cases found AMIE included the correct diagnosis 59% of the time versus 34% for unassisted clinicians.
- A meta-analysis in npj Digital Medicine pooling 50 studies across 25 models concluded AI systems often perform comparably to — and in some specialties better than — clinicians on standardized diagnostic and triage tasks.
Performance Under Uncertainty And Real-World Use
When tests mimic clinical uncertainty or unscripted interactions, performance falls. An NEJM AI study using a 750-question script concordance benchmark found even top models (e.g., OpenAI’s o3) achieved only about 68% accuracy — below senior residents and attending physicians. Similarly, hallucination studies have shown LLMs can accept and elaborate on fabricated clinical details 50–83% of the time under default settings.
A Stanford-led benchmark evaluating 20 models and four clinical AI tools across 1,100 cases scored potential harm if users directly applied the recommendations: direct application risked severe harm in 24.6% of cases, and more than 80% of those severe failures were omission errors — failing to flag critical dangers rather than inventing facts.
Why Deployment And Safeguards Matter
The evidence indicates two key points: (1) in narrow, structured tasks LLMs can match or exceed human performance; and (2) in open-ended, unsupervised interactions — the way many patients actually use chatbots — models are more likely to make dangerous omissions or confidently repeat false inputs. In short, capability in a test environment does not automatically translate into safe performance in the wild.
Design choices and guardrails — triage prompts, explicit refusal policies, escalation triggers, and integration with clinical workflows — determine whether a model’s high test scores translate into safer patient outcomes. The lawsuits now probe legal accountability when these systems are used by laypeople in crisis.
Legal, Ethical And Clinical Stakes
The Winters case and other pending suits will test how courts assign responsibility when users rely on chatbot output during medical or psychological emergencies. Regulators and health systems must weigh potential benefits — faster differential diagnoses, improved triage, better access to information — against the real risks of omission, hallucination, and inappropriate reassurance in unsupervised settings.
Bottom line: Research shows real diagnostic potential for AI under structured conditions, but substantial failure modes remain in unscripted, real-world use. The question for courts, hospitals, and developers is not only "How smart is the model?" but "How is it deployed, supervised and limited when people use it for health decisions?"
Originally published on Forbes.com.
Help us improve.

























