PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 4, 2025Menopause The Journal of The North American Menopause Society4 citations

Evaluation of the accuracy and readability of large language model responses on menopause and hormone therapy

View Full Paper
JKJana KaramCSChrisandra ShufeltNSNancy Safwan

Key Points

  • LLMs showed limited accuracy in answering menopause-related questions, with variability among models.
  • ChatGPT 3.5 exhibited the highest accuracy of 70% for patient-level queries, outperforming Gemini significantly.
  • Assessment using Flesch Reading Ease Score revealed difficulties in comprehension, especially with Gemini's lower score.
  • Improving LLM responses is crucial for providing reliable menopause and hormone therapy information for patients and clinicians.

Abstract

Objective: Generative artificial intelligence is rapidly evolving and is now being explored in health care to support patient and clinician education. This study evaluated the accuracy, completeness, and readability of four large language models (LLMs): ChatGPT 3.5, Gemini, ChatGPT 4.0, and OpenEvidence in answering questions about menopause and hormone therapy. Methods: A total of 35 questions (20 patient-level, 15 clinician-level) were entered into each LLM. OpenEvidence was only used for clinician-level questions. Four blinded expert reviewers rated responses as accurate and complete, accurate but incomplete, or inaccurate. Readability of patient-level responses was assessed using the Flesch Reading Ease Score (FRES) and word count. Analysis used ANOVA for readability, odds ratios for accuracy comparisons. Results: For patient-level questions, ChatGPT 3.5 achieved the highest accuracy (70%), followed by ChatGPT 4.0 (60%) and Gemini (30%); Gemini had significantly lower odds of accuracy compared with ChatGPT 3.5 (OR=0.18, 95% CI=0.05-0.71; P =0.014). FRES scores differed significantly ( P 0.05). Conclusion: LLMs demonstrated limited accuracy and frequent incorrect or incomplete responses to menopause-related queries, highlighting the need to improve model performance to ensure accurate and reliable information for both patients and clinicians.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Karam et al. (2025) studied this question.

synapsesocial.com/papers/6930e8dbea1aef094cca3de9https://doi.org/10.1097/gme.0000000000002695
Ask AI
Helpful
Bookmark
Share
View Full Paper