PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 22, 20242 citations

Enhancing Code-Mixing in Named Entity Recognition: A Comprehensive Survey of Deep Learning Models

View Full Paper
OKOmkar KhadeSJShruti JagdaleGTGauri Takalikar

Key Points

Key points are not available for this paper at this time.

Abstract

In today's interconnected and multilingual world, it is common to come across code-mixed text that combines multiple languages. That being said, Named Entity Recognition, a critical task for extracting meaningful information from code-mixed text, is made extremely difficult by this linguistic variation. It is critical to narrow the accessibility gap in language technology by expanding the use of code-mixed NER models to low-resource languages and dialects, which frequently lack extensive linguistic resources. The importance of semi-supervised and unsupervised methods for code-mixed NER is highlighted by addressing the lack of labeled data, especially in low-resource languages. Understanding fine-grained entity types is also necessary to improve the accuracy and usefulness of code-mixed NER models. Given the frequency of code-mixing in these contexts, it is imperative to adapt these models to informal language and various forms of code-switching in social media and user-generated material. Integrating with multimodal data analysis becomes essential for a thorough comprehension of code-mixing in a variety of contexts. Our suggested method initializes the NER model by using the code-mixed pre-trained HingBERT language model, which is trained on a dataset comprising text in both Hindi and English. A CRF (Conditional Random Fields) classifier is incorporated into the architecture to model sequence dependencies in NER, while a Bi-LSTM layer provides context-based information. This novel method creates new opportunities for Code-Mixed NER, which aims to identify named things in code-mixed text and categorize them. In our linguistically varied world, the integration of HingBERT, CRF, and Bi-LSTM contributes to effective multilingual entity identification, marking a substantial advancement in resolving the challenges of Code-Mixed NER.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Khade et al. (2024) studied this question.

synapsesocial.com/papers/68e781e8b6db6435876f4b54https://doi.org/10.1109/ic-etite58242.2024.10493709
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A Survey of Current Datasets for Code-Switching Research2020 · 98 citations
  2. 2HinGE: A Dataset for Generation and Evaluation of Code-Mixed Hinglish Text2021 · 36 citations
  3. 3Review of Research on Named Entity Recognition2022 · 8 citations
  4. 4Hate Speech Detection in Hindi-English Code-Mixed Social Media Text2019 · 107 citations
  5. 5Improvised Transformer Network for NER on Low Resource EnglishHindi Code-Mixed Language from Scratch2020 · 6 citations