Key points are not available for this paper at this time.
Abstract Automating bridge inspections requires more than detecting individual damage instances. It demands systems capable of describing, contextualizing, and interpreting damage in an inspection‐relevant manner. Conventional computer vision approaches, such as object detection and segmentation, primarily address visual recognition tasks and are constrained by predefined damage classes. Recent advances in vision‐language models (VLMs), particularly foundation models with few‐shot and zero‐shot learning capabilities, offer new opportunities by enabling semantically rich descriptions and reasoning over visual scenes. This paper presents an emerging‐trend review of VLM applications in bridge inspection, with a specific focus on their ability to satisfy information requirements derived from inspection practice. The synthesis identifies promising research directions, such as image captioning, damage cause inference, and human–AI collaboration, while also revealing that existing studies often leverage VLMs only superficially, frequently limiting them to classification or rudimentary captioning. Persistent challenges remain, including domain adaptation, data quality, image scale, and the absence of multimodal benchmarks. To address these limitations, the paper proposes a conceptual, inspection‐oriented VL benchmark framework and outlines research directions required to move beyond damage classification toward operational multimodal systems that support documentation and interpretation.
Çelik et al. (Tue,) studied this question.