PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 26, 2026Electronics0 citationsOpen Access

A Contrastive and Uncertainty–Aware Framework for Multimodal Named Entity Recognition

View Full Paper
XYXiao YangInstitute of State AdministrationRZRuixue ZhaoAgricultural Information InstituteHLHonglei LiHunan Police Academy

Key Points

  • This study aims to enhance multimodal named entity recognition by addressing issues in text-image alignment and entity representation.
  • Proposed a contrastive uncertainty-aware framework (CUA-MNER) for MNER.
  • Implemented hierarchical vision-text alignment for improved token, phrase, and sentence correspondences.
  • Used variational uncertainty-aware fusion to manage modality contributions and enhance entity recognition.
  • Achieved F1 scores of 76.97% and 89.66% on Twitter2015 and Twitter2017 benchmarks, respectively.
  • Outperformed competitive baselines by 0.66 and 1.95 F1 points.
  • Identified that the model's components provide complementary benefits, indicating robustness in multimodal recognition.

Abstract

Multimodal named entity recognition (MNER) aims to improve entity detection in social media texts by leveraging accompanying images, but its performance is often affected by weak text–image alignment, noisy or irrelevant visual content, and limited separation among entity representations. To address these issues, this study proposes CUA-MNER, a contrastive uncertainty–aware framework that combines hierarchical vision–text alignment, variational uncertainty–aware fusion, and token-level contrastive learning. The alignment module models correspondences at token, phrase, and sentence levels, allowing local visual regions and global image context to support textual entity recognition. The fusion module estimates epistemic and aleatoric uncertainty through variational inference and adaptively adjusts the contribution of each modality for different samples. The contrastive objective further encourages entity representations of the same type to be closer while separating different entity types. Experiments on the Twitter2015 and Twitter2017 benchmarks demonstrate that CUA-MNER achieves F1 scores of 76.97% and 89.66%, respectively, outperforming competitive baselines by 0.66 and 1.95 F1 points. Ablation and diagnostic analyses show that the three components provide complementary benefits. These results suggest that modeling modality reliability is useful for robust MNER, while the additional modules also introduce computational overhead and leave cross-domain generalization as an open issue.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2026) studied this question.

synapsesocial.com/papers/6a3e17d3030ad1a9b3091310https://doi.org/10.3390/electronics15132770
Ask AI
Helpful
Bookmark
Share
View Full Paper