Key points are not available for this paper at this time.
Introduction Artificial intelligence (AI) chatbots has been well studied in many common diseases. However, little was reported for chordoma, which is a rare disease with high rates of recurrence, disability, and mortality. Research question This study aims to assess the performance of oncologists and state-of-the-art AI chatbots in response to frequently-asked questions (FAQs) of chordoma from real world. The response performance was mainly evaluated by overall quality, empathy, and readability. Material and methods Sixty chordoma-related FAQs, collected from social media, were addressed by various chatbots and oncologists, and the best-performing chatbot-generated text was further edited and assessed again. Rated scores were ordered for quality and empathy in a blind way. The readability was measured objectively by calculating Flesch-Kincaid Grade Level (FKGL), Automated Readability Index (ARI), and Gunning-Fog Index (GFI). Results AI chatbots were universally superior to oncologists in response quality (3.86±0.14 vs. 3.12±0.25, p<0.001) and empathy (3.28±0.41 vs. 2.95±0.48, p<0.001). DeepSeek-R1 achieved highest rated score in response quality (4.20±0.22), while Claude 3.5 Sonnet was considered as the best chatbots by comprehensive assessments. The chatbot drafted responses were easier to understand from patient’s perspective (p<0.001). Improved response quality (4.09±0.12, p<0.001), empathy (4.00±0.39, p<0.001), and readability (FKGL: 11.30±2.42, p<0.001) were obtained after editing the Claude-3.5-generated responses by oncologists. Discussion and conclusions AI chatbots reached favorable quality and empathetic performance in response to chordoma-related FAQs, and generated equivalent readability compared to oncologists. With generative chatbot’s assistance, oncologists may response more comprehensively and efficiently in addressing chordoma patient’s common inquiries.
He et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: