Text-based and multimodal generative foundational models-in this article termed "large language models" (LLMs)-have the capability to process textual data and images due to their multimodal capabilities [1,2].There has been a remarkable surge in published studies that report on the accuracy of LLMs in medical applications, reflecting LLMs' potential to significantly reshape healthcare [3][4][5].These studies represent a new genre of medical research.However, the methodology and presentation of results in these studies are highly variable [6].Inconsistent and incomplete reporting hampers the ability of the reviewers and readers to evaluate the methodology and results of the studies, as well as to assess the replicability of the findings.Consequently, there is a pressing need for guidelines to improve the quality of research reports that present the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM)
No takes yet. Share an insight, caveat, or question.
Park et al. (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: