Surgical site infections (SSIs) are common postoperative complications that increase patient morbidity, hospital stays, and healthcare costs. Early detection and precise documentation are critical for timely intervention and improved outcomes. While patients often submit images of their wounds through a patient portal for remote monitoring, manual review is challenging due to image variability, high volume, and subjectivity, underscoring the need for automated assessment tools. In this paper, we present Vision Transformer for CAptioning of Incision Images (ViTCAI), a vision-language model designed for automated triage and captioning of postoperative incision images. Fine-tuned on a clinically annotated dataset, ViTCAI improves descriptive accuracy in identifying surgical incisions and SSIs. Our results show that ViTCAI provides consistent, detailed captions that can support clinical decision-making, reducing workload and enhancing diagnostic efficiency in postoperative care.
Lee et al. (Mon,) studied this question.