Perplexity achieved the highest composite quality score (8.2/9) for generating hypertension treatment plans among 12 large language models, while Dyna AI scored the lowest (3.7/9) (p<0.00001).
Do different large language models vary in the quality, accuracy, and safety of treatment plans generated for stage two hypertension?
While LLMs generally provide detailed and guideline-adherent management plans for stage two hypertension, there is significant variability in quality, highlighting the need for regulation regarding reliability and safety.
p-value: p=<0.00001
Background: The use of large language models (LLMs) in clinical practice, medical education, and by patients is increasing rapidly. It is essential to ensure that information provided by these artificial intelligence (AI) chatbots is accurate and safe. Our goal was to analyze and compare hypertension treatment plans generated by popular LLMs and identify their strengths and limitations. Methods: ChatGPT-4o, Claude, ClinicalKey AI, Copilot (Microsoft), DeepSeek-V3, Dyna AI, Google Gemini, Grok, Meta AI, OpenEvidence, Perplexity, and Pi were prompted to generate a treatment plan for stage two hypertension. Five blinded reviewers scored each response in three domains: adherence to clinical guidelines, detail/clarity, and reliability/safety (sources provided/emphasis on seeing a healthcare professional). Mean scores for each domain were calculated and summed for a composite score. The responses were also analyzed qualitatively by three reviewers. Results: Perplexity received the highest composite score (8.2 out of 9), followed by OpenEvidence (7.7 out of 9). Dyna AI had the lowest overall score (3.7 out of 9), followed by Pi (4.8 out of 9), ClinicalKey AI (4.9 out of 9), and Meta AI (4.9 out of 9). Perplexity (3 out of 3), Grok (2.8 out of 3), and OpenEvidence (2.7 out of 3) had the highest scores for detailed and clear responses, while DynaAI had the lowest for both detail/clarity (1 out of 3) and reliability/safety (1 out of 3). ChatGPT-4o had the highest score for adherence to guidelines (2.7 out of 3) while Pi had the lowest (1.5 out of 3). Analysis of Variance (ANOVA) statistical test showed statistically significant differences across every subscore domain and composite scores (p=0.00125 for adherence to guidelines, p<0.00001 for detail/clarity, p=0.00003 for sources/seeing a professional, p<0.00001 for composite scores). Qualitatively, the LLMs tended to adhere to guidelines and provide sufficiently detailed management plans, but often did not provide sources and/or advise users to see a healthcare professional. Conclusions: The LLMs nearly always provided detailed management plans for stage two hypertension that adhere to clinical guidelines. However, there was significant variability in quality by different chatbots. Notably, medicine-specific LLMs were not necessarily superior to LLMs used by the general public. AI chatbots may require greater regulation to ensure that they inform users to see a healthcare professional and provide reliable sources.
Metzger et al. (Tue,) conducted a other in Stage two hypertension. Large Language Models (LLMs) vs. Comparative analysis among 12 LLMs was evaluated on Composite score of adherence to clinical guidelines, detail/clarity, and reliability/safety (out of 9) (p=<0.00001). Perplexity achieved the highest composite quality score (8.2/9) for generating hypertension treatment plans among 12 large language models, while Dyna AI scored the lowest (3.7/9) (p<0.00001).