Non-autoregressive methods enhance image captioning, improving accuracy and reducing hallucinations in generated captions.
While autoregressive models have achieved remarkable success in image captioning, their slow inference speed limits their applicability in real-time scenarios. Non-autoregressive methods provide a promising alternative for faster caption generation; however, they still encounter significant challenges. In particular, they struggle to capture complex content and abstract concepts necessary for producing semantically rich and accurate captions, which hinders the bridging of the image–text gap. Moreover, the generation process often leads to object hallucination–instances where incorrect or nonexistent objects are described, resulting in captions that misalign with the actual visual content. To address these issues, we propose the Vision-Text Semantic Reconstruction and Contrast (VTSRC) mechanism, which consists of two key modules. The first is the Visual-Text Reconstruction Network (VRN) module, which reconstructs visual representations into textual space, enriching captions with contributive and complex semantics to bridge the image-text gap. The second is the Visual Contrastive Generation (VCG) module, which leverages visual uncertainty to contrast distributions, recalibrating the model’s output and significantly reducing the incidence of hallucination, thereby generating coherent linguistic representations. Extensive evaluations demonstrate that our approach markedly improves the creation of semantically rich image captions, considerably reducing the frequency of hallucinations while maintaining high descriptive accuracy. Experimental results demonstrate that VTSRC achieves competitive performance on the challenging MSCOCO image captioning dataset, reaching the best CIDEr score of 133.9% on the COCO-caption Karpathy split to date.
No takes yet. Share an insight, caveat, or question.
Wang et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: