PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 22, 20260 citationsOpen Access

Challenges in Named Entity Recognition for a Morphologically Rich Marathi Language

View Full Paper
MRMr. Sonar Nandkishor RajendraMVMs. Jagtap Aparna Vijay

Key Points

  • The aim is to explore challenges in named entity recognition for the Marathi language, focusing on linguistic and computational issues.
  • Discussed linguistic complexities of Marathi including morphology and inflection.
  • Examined the impact of data scarcity and domain variation on NER performance.
  • Analyzed the role of corpus development for enhancing NER systems.
  • Considered the application of deep learning techniques for improving outcomes.
  • Identified significant challenges due to Marathi's rich morphology and flexible word order.
  • Established that data scarcity and lack of standardized tools hinder NER performance.
  • Proposed language-specific modelling as crucial for developing effective NER systems.

Abstract

Named Entity Recognition (NER) is an essential task in Natural Language Processing (NLP) that focuses on identifying and classifying proper names such as persons, places, organizations, dates, and other meaningful entities within textual data. Although NER systems have achieved remarkable success for widely studied languages like English, their effectiveness for Indian languages remains limited. Marathi, a prominent Indo-Aryan language written in the Devanagari script, presents unique linguistic complexities including rich morphology, extensive inflection, flexible word order, and the absence of capitalization. These characteristics, along with the lack of large annotated datasets and standardized tools, make the task of Named Entity Recognition particularly challenging. This paper presents a comprehensive discussion of the linguistic and computational issues encountered while developing NER systems for Marathi. It examines the impact of morphological variation, lexical ambiguity, orthographic inconsistencies, data scarcity, and domain variation on NER performance. The study concludes by emphasizing the importance of language-specific modelling, corpus development, and the adoption of advanced deep learning techniques for improving Marathi NER systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rajendra et al. (2026) studied this question.

synapsesocial.com/papers/69bf390ac7b3c90b18b431aahttps://doi.org/10.5281/zenodo.19124976
Ask AI
Helpful
Bookmark
Share
View Full Paper