Key points are not available for this paper at this time.
Large language models (LLMs) such as ChatGPT have demonstrated remarkable performance, including passing professional exams. However, because they generate responses through probabilistic prediction, their ability to directly replace medical experts remains limited. This study evaluates the applicability of LLMs in medicine using models available as of August 2023. Two medical guidelines were selected, and key questions derived from them were used to assess three offline models (KoVicuna, WizardVicuna, and LLaMa2) and the online ChatGPT model via LangChain. Model performance was evaluated based on accuracy and response time. ChatGPT achieved the highest accuracy with the shortest response time. Among the offline models, WizardVicuna 13 B exhibited high accuracy, whereas LLaMa2 7 B demonstrated balanced performance with relatively fast responses. Although LLMs cannot provide precise diagnoses or treatment recommendations owing to hallucinations and computational constraints, they show promise as clinical decision-support tools. With further refinement, LLMs may augment rather than replace physicians in medical practice.
Sim et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: