CRBC News
Health

Study Finds AI Makes Kidney-Recipient Choices With Unwarranted Certainty

Study Finds AI Makes Kidney-Recipient Choices With Unwarranted Certainty
Photo Credit: iStock

A Penn State study presented at ACM FAccT 2026 found that large language models often make confident kidney-allocation decisions where humans would hesitate. The models tended to overweight single traits (for example, alcohol use) and rarely chose a 'flip-a-coin' or uncertain option. Researchers warn this false precision could be dangerous in medical and other high-stakes settings and call for stronger guardrails and models that surface uncertainty.

A new Penn State study presented at the 2026 ACM Fairness, Accountability and Transparency (FAccT) conference warns that large language models (LLMs) often express unwarranted certainty when asked to make ethically fraught medical decisions such as allocating scarce donor kidneys.

Researchers tested several prominent chatbots using fictional kidney-transplant scenarios that included patient details like age, overall health, alcohol use and number of dependents. They compared model answers to results from earlier studies of lay human participants (people not identified as medically trained).

How The Study Worked

The team varied the prompts to probe different decision strategies. Sometimes they isolated a single trait, sometimes they combined multiple factors, and in other cases they added an explicit "flip-a-coin" option to capture indecision — a common feature of real human moral judgment.

"Sometimes we isolated just one trait at a time. Sometimes we mixed several traits together to see how AI weighed competing factors, and sometimes we added a flip-a-coin option to measure indecision, a key factor present in human moral judgment," said Hadi Hosseini, lead author and associate professor at Penn State.

Key Findings

Two consistent patterns emerged:

  • Overweighting Single Traits: Models often fixated on a single attribute (for example, alcohol use) rather than balancing multiple considerations the way people tend to do.
  • Rarely Expressing Uncertainty: Even when explicitly given an option to leave the decision to chance, the LLMs usually opted for a deterministic choice instead of endorsing indecision.

"AI models almost never do this: even when directly given the option to 'flip a coin,' they overwhelmingly commit to a confident, deterministic answer instead," Hosseini said. "That's a meaningful gap, since real moral dilemmas often don't have one clearly correct answer."

Why This Matters

Organ allocation is a high-stakes example where no perfectly objective answer exists. When AI presents a polished, confident verdict, it risks giving the impression that the decision is more objective than it is, potentially shaping clinical choices and policy in ways that understate nuance and ethical complexity.

The researchers emphasize they are not suggesting AI should replace clinicians. Rather, the findings highlight the need for robust guardrails, clearer communication of uncertainty, and design choices that let models acknowledge ambiguity instead of masking it with false precision.

"When we allocate something scarce, whether it's a kidney, a job, or access to some other resource, there isn't always a single objectively correct answer," said John Dickerson, CEO of Mozilla.ai and a collaborator on the study. "Humans recognize that ambiguity and codify it via open debate into the allocative process. AI models often don't."

Broader Implications

The concerns raised by this study extend beyond transplantation to other medical and high-stakes domains — from health chatbots and therapy tools to insurance adjudication — where hidden model certainty could have real-world consequences. The authors recommend developers build systems that surface uncertainty, and urge policymakers and health systems to require safeguards before adopting such tools for ethically sensitive decisions.

Data Availability: The study compared LLM responses to earlier human-subject results drawn from non-medical participants; the paper and presentation at ACM FAccT 2026 provide full methodological details.

Help us improve.

Related Articles

Trending