PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 21, 2026Digital Health2 citationsOpen Access

The reliability and readability of large language models in answering patient questions on maintenance hemodialysis: A comparative study

View Full Paper
JCJinjin CaoZWZhonghua WuZWZhishui Wu

Key Points

  • This research aims to evaluate the reliability and readability of information provided by five leading large language models (LLMs) related to maintenance hemodialysis (MHD).
  • Cross-sectional comparative design
  • Identification of seventeen MHD-related questions using Google Trends and online forums
  • Assessment of model responses with DISCERN, EQIP, JAMA, and GQS for reliability, along with readability indices like FKGL, FRES, and SMOG
  • Heatmap analysis for intra-model response variability
  • High inter-rater reliability confirmed between experts (ICC 0.851 to 0.879, P < .001)
  • Significant differences in reliability and readability among the five LLMs
  • Perplexity consistently achieved higher reliability scores than others (P < .001)
  • All models produced texts above the sixth-grade reading level, with FRES scores below the recommended range

Abstract

Objective This study aimed to systematically evaluate five leading LLMs—ChatGPT, DeepSeek, Copilot, Gemini, and Perplexity—in providing MHD-related health information. The primary objectives were to determine (1) the reliability of MHD-related information generated by LLMs and (2) whether its readability meets the recommended standards for patient educational materials. Methods A cross-sectional comparative design was adopted. The approximate timeframe during which the responses were generated was October 2025. Seventeen frequently asked MHD-related questions were identified using Google Trends and two online patient–caregiver forums. Each query was input into the five LLMs (ChatGPT-4o, Copilot, Gemini 2.5 Pro, Perplexity Pro, and DeepSeek-V3.2-Exp), and their responses were assessed using DISCERN, EQIP, JAMA, and GQS criteria for reliability, alongside FKGL, FRES, SMOG, CLI, ARI, and LWF readability indices. A heatmap analysis was also conducted to evaluate intra-model response variability. Results High inter-rater reliability was confirmed between the two experts (ICC for average measures ranged from 0.851 to 0.879, all P < .001). Significant differences were observed among the five LLMs in both reliability and readability. Overall reliability scores were relatively low; however, Perplexity consistently achieved higher DISCERN, EQIP, and JAMA scores compared with Gemini, ChatGPT, Copilot, and DeepSeek ( P < .001). In terms of readability, all models produced texts exceeding the sixth-grade reading level. Their ARI, GFI, FKGL, CLI, and SMOG scores were notably higher than recommended, while FRES scores were substantially below the 80–90 range. Heatmap analysis further demonstrated that although Perplexity and ChatGPT maintained relatively stable mean scores, they exhibited higher variability across different queries. Conclusions Current large language models (LLMs) exhibit significant variability in delivering maintenance hemodialysis information. While all five evaluated models demonstrated limitations in information quality, transparency, and readability, Perplexity performed relatively better overall. However, persistent deficiencies in source attribution, language accessibility, and response consistency limit their immediate clinical and educational utility. Future LLM development should prioritize readability optimization and context-aware customization to better support patient education.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cao et al. (2026) studied this question.

synapsesocial.com/papers/69be362d6e48c4981c674e56https://doi.org/10.1177/20552076261435836
Ask AI
Helpful
Bookmark
Share
View Full Paper