Key points are not available for this paper at this time.
Background: This study aims to compare the quality, reliability, and readability of information provided by artificial intelligence-based language models, ChatGPT-5 and DeepSeek V3, regarding foot and ankle disorders. Methods: The quality, reliability, and readability of the texts generated by both AI models were analyzed using DISCERN, the Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P), the Global Quality Score (GQS), and the CLEAR scoring system. DISCERN was used to assess information reliability, PEMAT-P to evaluate understandability and actionability, GQS to assess overall quality, and CLEAR to evaluate content quality and accuracy. Standardized questions were asked to both models for 35 different foot and ankle disorders, and the generated texts were evaluated by two independent orthopedic specialists using a blinded method. Readability analysis was performed using word count, the Flesch–Kincaid Grade Level (FKGL; required reading level), and the Flesch Reading Ease (FRE; ease of readability) scoring systems. Results: ChatGPT-5 scored significantly higher than DeepSeek V3 in DISCERN, PEMAT-P, GQS, and CLEAR evaluations (p < 0.05), indicating that ChatGPT-5 provides more reliable, comprehensive, and higher-quality information. DeepSeek V3 demonstrated better readability, producing simpler and more understandable content, as reflected in its lower FKGL score and higher FRE score. Conclusions: While ChatGPT-5 delivers more detailed and reliable health information, DeepSeek V3 offers simpler and more readable texts. Both models have distinct advantages for patient education. Future research should assess the impact of AI-generated health information on patient decision-making and its clinical application potential.
Karagoz et al. (Thu,) studied this question.