The increasing demand for medical imaging has significantly challenged healthcare systems, emphasizing the need for efficient and accurate diagnostic support tools. Recent advancements in computer vision (CV) and natural language processing (NLP) have demonstrated great promise in addressing these challenges, particularly through the automation of medical report generation. Automatic medical report generation (AMRG) has become a pivotal application of artificial intelligence (AI) in the medical domain, which involves extracting critical information from medical images and generating textual reports. These reports aid clinicians in analyzing image content more efficiently and accurately, thereby improving diagnostic precision. This article provides a comprehensive review of recent advancements in AMRG, with a particular focus on the commonly—employed methodologies, including convolutional neural network (CNN), recurrent neural network (RNN), Transformers and their variants, and large language model (LLM) into AMRG. Moreover, this review also examines both widely—used and less frequently-used datasets and compares various evaluation metrics to provide an in-depth analysis of different AMRG methodologies. Finally, key achievements and future research directions in the field are summarized, highlighting challenges such as cross-modal fusion, model interpretability, and data privacy protection, while suggesting potential future trends in the development of this technology.
Yan et al. (Fri,) studied this question.