PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 28, 2026Digital Health1 citationsOpen Access

Alignment of large language models in solving medical ethical dilemmas

View Full Paper
VSVera SorinBGBenjamin S. GlicksbergPKPanagiotis Korfiatis

Key Points

  • The aim is to examine how large language models reconcile moral reasoning in medical ethical dilemmas.
  • Tested 14 large language models using five medically adapted Trolley dilemmas.
  • Each model generated 10 independent answers per dilemma, totaling 700 responses.
  • Responses were categorized as utilitarian ('yes') or deontological ('no').
  • χ2 tests were used to compare the models' choices.
  • Utilitarian responses varied widely, with some models selecting this option in only 20% of cases.
  • Models like Qwen-2-7B favored utilitarian choices in 80% of scenarios (p < .001).
  • Some models recommended impermissible actions, including nonconsensual limb amputation.
  • 44 out of 700 outputs (6.3%) endorsed actions violating ethical norms.

Abstract

Background Large language models (LLMs) are entering clinical workflows, yet their alignment with medical-ethical principles is unclear. Objective To evaluate how LLMs reconcile deontological and utilitarian reasoning in medically adapted “Trolley dilemma” scenarios. Methods We tested 14 LLMs, including GPT-o1-preview and DeepSeek-R1-Distill-Llama, using five medically adapted Trolley dilemmas sourced from philosophical and real-world scenarios. Each model generated 10 independent answers per case (overall 700 responses). Responses were forced to “yes” (utilitarian) or “no” (deontological). The primary outcome was the proportion of utilitarian choices. χ 2 tests compared models. Results LLM responses varied. Some models, such as GPT-4 and Gemma, selected the utilitarian option in 20% of cases, while others, like Qwen-2-7B, reached 80% ( p < .001). GPT-o1-preview and DeepSeek-R1-Distill-Llama selected utilitarian responses in 44% and 38% of cases, respectively. Some models selected the consequence-maximizing option in vignettes where doing so required overriding consent or endorsing intentional harm within the scenario framing. Seven of the fourteen models tested recommended impermissible medical actions, including nonconsensual limb amputation, killing a healthy person for organ harvesting, and transferring a patient to hospice against their wishes, in 44 of 700 outputs (6.3 %). Conclusions In this constrained forced-choice stress test, models occasionally endorsed boundary-violating actions within the vignette framing. These results do not generalize to routine clinical decision-making, but motivate further benchmarked evaluation of refusal behavior and boundary constraints in ethically sensitive medical contexts. We must carefully define how these models influence clinical decisions, establish robust ethical guidelines for their use, and ensure that fundamental ethical norms are not compromised.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sorin et al. (2026) studied this question.

synapsesocial.com/papers/69a287010a974eb0d3c02508https://doi.org/10.1177/20552076261428395
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Ethical Alignment of LLMs in Healthcare: Does GPT-o1 Adopt a Deontological or Utilitarian Approach?2024 · 2 citations
  2. 2Corpus-Based Evaluation of Decision-Making in Medical Ethics by Large Language Models2025
  3. 3GALATEA II: Benchmarking LLM Safety in Clinical Simulation. Behavioural Safety and Ethical Robustness of Large Language Models in a Multi-Agent ICU Decision Support Architecture2026
  4. 4Large Language Models Amplify Human Biases in Moral Decision-Making2024 · 4 citations
  5. 5Algorithmic bias in clinical resource allocation by a large language model: a cross-sectional in-silico evaluation of 13,608 decisions2026