Objectives: To assess and check the validity and reliability of the answers given by ChatGpt-4o mini, Deepseek, Copilot and Gemini 1.5 flash daily chatbots to often seeked queries in the area of periodontology. Materials and Methods: Questions were selected from the most frequently asked patient questions by a periodontologist. Each question was asked to the chatbots three times. The answers (n=240) were independently evaluated by two periodontologists on a Likert scale (5=violently agree; 4=agree; 3: neutral; 2=disagree; 1=violently disagree). Disputes in scoring were removed through evidence-based negotiations. In evaluating the validity of the answers: Low threshold was determined as a score ≥4 for whole three answers; high threshold was determined as a score 5 for whole three answers. Fisher's exact test was performed to compare the validity of the answers among the chatbots. Cronbach's alpha was computed to evaluate the consistency and reliability of recurrent answers for each chatbot. Results: All four chatbots answered the questions. In the low-threshold validity test, ChatGpt had 100%, Deepseek and Copilot had 95%, Gemini had 65%. Gemini was significantly different from the others (p0.05), both were significantly higher than Copilot and Gemini (p
Mahmure Ayşe Tayman (2025) studied this question.