PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 18, 2026Health care science0 citationsOpen Access

Evaluating ChatGPT's Adherence to Medical Ethics: A Prerequisite for Artificial Intelligence in Medicine

View Full Paper
YZYing ZhangChinese Academy of Medical Sciences & Peking Union Medical CollegeYLYuyang LiuChinese Academy of Medical Sciences & Peking Union Medical CollegeTLTingyu LvChinese Academy of Medical Sciences & Peking Union Medical College

Key Points

  • This research evaluates ChatGPT's performance in addressing medical ethics questions compared to human experts.
  • Developed a Medical Ethics Evaluation dataset with 465 single-choice questions from medical ethics standards.
  • Assessed two AI models, GPT-3.5 and GPT-4, against responses from medical ethics experts.
  • Performed independent tests to ensure consistency and calculated accuracy for each model and expert.
  • Used chi-square tests to compare the performance of AI models with that of human experts.
  • GPT-3.5 achieved 38.92% accuracy; GPT-4 achieved 27.10%.
  • Two medical ethics experts had accuracies of 86.23% and 78.32%.
  • Both experts significantly outperformed AI models, indicating a large gap in ethical understanding.
  • AI models showed some alignment with core medical ethics principles, especially in ethical dilemmas.

Abstract

ABSTRACT Background As artificial intelligence continues to play an expanding role in healthcare, ensuring its compliance with medical ethics is essential. However, the ethical performance of artificial intelligence in medical contexts remains insufficiently studied. This study aimed to evaluate the ability of ChatGPT to address questions related to medical ethics and to compare its performance with that of human experts. Methods A Medical Ethics Evaluation dataset was developed, consisting of 465 single‐choice questions derived from a range of medical ethics standards. These questions were used to assess two artificial intelligence models, GPT‐3.5 and GPT‐4. Model responses were compared with those provided by two medical ethics experts. Each test was conducted independently twice to ensure consistency. Accuracy was calculated for each model and expert, and chi‐square tests were used to compare differences in performance. Results GPT‐3.5 achieved an overall accuracy of 38.92%, while GPT‐4 achieved 27.10%. In comparison, two medical ethics experts achieved substantially higher accuracies of 86.23% and 78.32%, respectively. Both experts performed significantly better than GPT‐3.5 and GPT‐4. These findings indicate a substantial gap between artificial intelligence models and human experts in understanding and applying medical ethics principles. The relatively low performance of the models, compared with their reported strengths in diagnostic tasks, may reflect the complexity and nuance of ethical reasoning in medicine. Nevertheless, the large language models showed some ability to align with core medical ethics principles, particularly in ethical dilemma scenarios, and were also able to generate responses that addressed psychological needs. Conclusions Artificial intelligence models currently show limited accuracy in medical ethics decision‐making compared with human experts. Although these models demonstrate some alignment with fundamental ethical principles, the performance is not yet sufficient for reliable use in ethically sensitive medical contexts. Further optimization is needed to improve their ability to meet the ethical demands of medical practice.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/69e320af40886becb653fd34https://doi.org/10.1002/hcs2.70067
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Assessing Generative Pretrained Transformers (GPT) in Clinical Decision-Making: Comparative Analysis of GPT-3.5 and GPT-42024 · 58 citations
  2. 2Evaluating ChatGPT’s moral competence in health care-related ethical problems2024 · 2 citations
  3. 3Evaluating the Potential and Accuracy of ChatGPT-3.5 and 4.0 in Medical Licensing and In-Training Examinations: Systematic Review and Meta-Analysis2025
  4. 4Assessing the performance of ChatGPT in addressing ethical dilemmas in oncology.2026
  5. 5Evaluating the Potential and Accuracy of ChatGPT-3.5 and 4.0 in Medical Licensing and In-Training Examinations: Systematic Review and Meta-Analysis (Preprint)2024