PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 18, 2025PLoS ONE2 citationsOpen Access

Evaluating the quality of ChatGPT-generated medical information on major ophthalmic conditions: A comparative assessment against the EQIP tool and guidelines

View Full Paper
MHM. W. HuPZPingping ZouLTLi Teng

Key Points

  • ChatGPT's responses achieved an EQIP score of 18, indicating quality in medical content generation.
  • Inter-rater reliability was high, with a Cohen's kappa value of 0.926, showcasing strong agreement between evaluators.
  • The AI's alignment with established clinical guidelines was noted at 84%, suggesting effective information delivery.
  • Further enhancements are necessary to improve the precision of quantitative medical data provided by ChatGPT.

Abstract

Background: The use of artificial intelligence for creating medical information is on the rise. Nonetheless, the accuracy and reliability of such information require thorough assessment. As a language model capable of generating text, ChatGPT needs a detailed examination of its effectiveness in the healthcare domain. Objective: This research sought to evaluate the precision of medical data produced by ChatGPT-4o ( https://chat.openai.com/chat , accessed Mar. 12, 2025), concentrating on its capability to handle the top five ophthalmic issues that pose the greatest global health challenges. Furthermore, the investigation compared the AI’s answers to recognized medical guides. Methods: This research employed an adapted version of the Ensuring Quality of Information for Patients (EQIP) instrument to evaluate the quality of ChatGPT’s replies. The guidelines for the five conditions were rephrased into pertinent queries. These queries were then fed into ChatGPT, employing benchmarking against established ophthalmology clinical guidelines, and the resulting answers were independently scrutinized for precision and consistency by two investigators. The consistency among raters was evaluated using Cohen’s kappa value. Results: The median EQIP score across the five conditions was 18 (IQR 18-19). The modified EQIP instrument revealed a robust consensus between the two evaluators when assessing ChatGPT’s responses, as indicated by a Cohen’s kappa value of 0.926 (95% CI 0.875-0.977, P<0.001). The alignment between the ChatGPT responses and the guideline recommendations was 84% (21/25), as indicated by a Cohen’s kappa value of 0.658 (95% CI 0.317-0.999, P<0.001). Conclusions: ChatGPT demonstrates robust quality and guideline compliance in producing medical content. Nevertheless, improvements are necessary to enhance the accuracy of quantitative data and ensure a more comprehensive coverage, thereby offering valuable insights for the advancement of medical information generation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hu et al. (2025) studied this question.

synapsesocial.com/papers/68f3793258f37cefb60d34d1https://doi.org/10.1371/journal.pone.0334250
Ask AI
Helpful
Bookmark
Share
View Full Paper