PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 10, 2025Scientific Reports4 citationsOpen Access

Medical triage as an AI ethics benchmark

View Full Paper
NKN. KirchKHKonstantin HebenstreitMSMatthias Samwald

Key Points

  • Most large language models demonstrate improved ethical decision-making in medical dilemmas, outperforming random guessing.
  • Open-source models tend to make more serious ethical errors than proprietary models during ethical evaluations.
  • Guiding ethical principles hinder the performance of AI models in triage, contrary to findings from other machine ethics benchmarks.
  • Adversarial prompts lead to significant decreases in model accuracy, emphasizing the importance of context in ethical AI evaluations.

Abstract

We present the TRIAGE benchmark, a novel machine ethics benchmark designed to evaluate the ethical decision-making abilities of large language models (LLMs) in mass casualty scenarios. TRIAGE uses medical dilemmas created by healthcare professionals to evaluate the ethical decision-making of AI systems in real-world, high-stakes scenarios. We evaluated six major LLMs on TRIAGE, examining how different ethical and adversarial prompts influence model behavior. Our results show that most models consistently outperformed random guessing, with open source models making more serious ethical errors than proprietary models. Providing guiding ethical principles to LLMs degraded performance on TRIAGE, which stand in contrast to results from other machine ethics benchmarks where explicating ethical principles improved results. Adversarial prompts significantly decreased accuracy. By demonstrating the influence of context and ethical framing on the performance of LLMs, we provide critical insights into the current capabilities and limitations of AI in high-stakes ethical decision making in medicine.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kirch et al. (2025) studied this question.

synapsesocial.com/papers/68c1ce7054b1d3bfb60f5bd5https://doi.org/10.1038/s41598-025-16716-9
Ask AI
Helpful
Bookmark
Share
View Full Paper