Randomized trial investigates if image descriptions enhance machine translation evaluation metrics sensitivity, indicating the importance of visual context.
Multimodal machine translation (MMT) aims to integrate visual context with textual data to improve the translation of ambiguous source text, such as the inclusion of an image as additional context. However, the evaluation of systems largely still relies on automatic metrics designed to evaluate text alone, and do not account for additional modalities during evaluation. The lack of dedicated MMT evaluation methods often results in inconsistent findings and creates uncertainty regarding the actual contribution of visual context in translation. In this work, we examine the performance of state-of-the-art trained and untrained evaluation metrics, particularly when comparing multimodal and text-only systems. Our evaluation focuses on the degree to which existing metrics are sensitive enough to distinguish between multimodal and text-only machine translation systems. We further investigate the potential for automatically generated image descriptions to serve as effective contextual signals for improving metric sensitivity to multimodal tasks. Our results show that incorporating such visual information into supervised metrics yields better alignment with human judgment. While all metrics successfully distinguished image-aware from image-agnostic systems on general test sets, both n-gram-based and embedding-based metrics struggled with respect to contrastive evaluation designed to capture context-dependent errors. Furthermore, we discuss how the presence of visual context may influence human evaluator judgment, observing that given the opportunity, human ratings are often substantially revised, further emphasizing the critical role of context in the evaluation of MMT.
No takes yet. Share an insight, caveat, or question.
Haq et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: