Key points are not available for this paper at this time.
Background/Objectives: Large language models (LLMs) are increasingly consulted for information about cleft lip and palate (CLP), yet the reliability of their outputs across clinical domains has not been evaluated. This study aimed to compare the quality of CLP-related information generated by GPT-4o and Gemini 2.5 Pro across multiple thematic domains using a validated quality instrument and a reliability-first analytic framework. Methods: Fifty-four standardized CLP questions across six domains were submitted to GPT-4o (OpenAI) and Gemini 2.5 Pro (Google DeepMind) on 25 September 2024 via their public interfaces, using new, history-free sessions and default settings, yielding 108 responses. Three independent, CLP-experienced raters scored each response using the Global Quality Score (GQS; 1–5 scale assessing accuracy, completeness, and clinical usefulness). Before comparing models, we applied a reliability-first filter: only domains where all three raters showed substantial agreement (Fleiss’ kappa κ ≥ 0.60) were included in statistical comparisons. Domains that failed this threshold were analyzed qualitatively to identify the source of disagreement. A descriptive taxonomy of errors was developed for low-scoring responses. Results: Three domains met the reliability threshold (General Care Information, General Cleft Information, and Pre-Treatment Information; 30 paired questions). Both models performed at a high and practically equivalent level: GPT-4o median GQS 4.33 (IQR 4.00–5.00) versus Gemini 2.5 Pro 5.00 (IQR 4.00–5.00); the difference was not statistically significant (Wilcoxon V = 139.00, p = 0.691; Hodges–Lehmann median difference 0.00, 95% CI −0.33 to 0.67). Three domains were excluded because rater agreement was insufficient; qualitative review showed this reflected genuine clinical practice variation rather than clear model errors. The most common inaccuracies were overgeneralization of outcomes, outdated surgical timing, and omission of multidisciplinary team roles. Conclusions: Both models provided high-quality CLP information in domains supported by clinical consensus, indicating they may serve as useful adjuncts for general patient and family counseling. Clinicians should, however, verify any treatment-specific content against current institutional protocols before relaying it to patients. Future research should assess readability, alignment with health literacy, and patient comprehension of AI-generated CLP information.
Bilder et al. (Mon,) studied this question.