Chatbots are just as good as professors in both factual recall and clinical scenario analysis: Emergence of a new tool in clinical microbiology and infectious disease Dear Editors, We read with interest an article published by Theodosiou and Read on the recent and potential future applications of artificial intelligence and its relevance to clinical infection practice.1 Here, we describe the performance of five chatbots, namely ChatGPT, Perplexity, Claude, Copilot and Gemini; as well as the free and subscribed versions of ChatGPT, Perplexity and Claude; in answering multiple choice questions (MCQs) and true/false (T/F) questions in clinical microbiology and infectious disease.Their performances were also compared with those of professors and medical students.Three books with clinical microbiology and infectious disease questions were used: Book 1, an e-book of T/F questions 2 ; Book 2, an e-book of MCQs 3 ; and Book 3, a physical book of MCQs. 4 All questions requiring interpretation and/or analysis of image were excluded.As a result, a total of 2637 questions were used.Each question was used for testing the five chatbots, two professors of clinical microbiology and infectious disease, and three final-year medical students.Answers provided by the books were considered as the correct answers.One mark was awarded for every correct answer.No mark was deducted for wrong answers.For T/F questions, 0.5 marks were awarded for a pass.For MCQs, 0.2 marks were awarded for a pass.To estimate the speed of performance, the first 100 questions of Book 2 were grouped into 10 sets, whereas the first 25 questions of Book 3 were grouped into five sets, and each question set was fed to the chatbots.The runtime is defined as the time between the pressing of the "Enter" key on the chatbot and appearance of results.Overall, there was no significant difference among the median scores obtained by the five free chatbots for the questions in the three books.Three chatbots (ChatGPT, Perplexity and Claude) are available in both free (ChatGPT 3.5, Perplexity and Claude Sonnet) and subscribed (ChatGPT 4.0, Perplexity Pro and Claude Opus) versions.For the factual recall questions (Book 1 and 2), there was no significant difference between the median scores obtained by the free (85%) and subscribed (87%) versions, but for the clinical scenario questions (Book 3), the median score obtained by the subscribed Journal of Infection 89 (2024) 106274
No takes yet. Share an insight, caveat, or question.
Wang et al. (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: