This study comparatively evaluated accuracy and response stability of artificial intelligence (AI)-models in answering diagnostic, therapeutic and prognostic questions related to dental trauma (DT), based on International Association of Dental Traumatology (IADT) guidelines. Fifty multiple-choice questions derived from IADT guidelines were categorised into diagnostic (n = 13), therapeutic (n = 30) and prognostic (n = 7) domains and administered to eight AI-models once weekly over three consecutive weeks. Responses were coded as correct (1) or incorrect (0) and analysed. Statistical significance was set at p 0.05). Model type significantly affected accuracy (p = 0.001), whereas question category (p = 0.259) and time (p = 0.436) had no effect. AI models showed heterogeneous performance. High accuracy did not necessarily correspond to response stability, as observed for ChatGPT-4.5, indicating that these systems should be used cautiously and only as supplementary tools within a structured multiple-choice framework.
Kilic et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: