Multimodal named entity recognition improves prediction using prior knowledge and text-guided integration, suggesting enhanced accuracy in entity extraction.
Key Points
The aim is to enhance Multimodal Named Entity Recognition by integrating external knowledge and improving cross-modal fusion.
Developed a framework called PKTF with two main stages: prior assisted knowledge generation and entity recognition.
Utilized Intern VL2-8B to generate contextual prior knowledge for original text.
Designed Text-Max-Directed Fusion Module to focus attention scores based on text guidance.
Achieved F1-scores of 75.43% on the Twitter-2015 dataset and 88.74% on the Twitter-2017 dataset.
Demonstrated competitive performance compared to existing multimodal models.