PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 2, 20240 citations

KULemma: Towards a Comprehensive Bangla Lemmatizer

View Full Paper
SISaimum IslamKAKazi Masudul Alam

Key Points

Key points are not available for this paper at this time.

Abstract

Lemmatization is a natural language processing (NLP) based text normalization technique that effectively improves data consistency and assists in context interpretation. Nevertheless, due to the highly inflected nature and morphological richness of linguistics, lemmatization in Bangla NLP holds a thorny challenge. In this research work, we take the challenge to build a comprehensive Bangla lemmatizer. We concreted homogeneous resources by compiling heterogeneous resources, assembled root words, collected linguistic rules, and applied the Trie along with "Longest Substring Search by Removing Affix (LSSRA)" to develop the lemmatizer. Our system aimed to lemmatize words based on their parts of speech within a given sentence and utilized the sequences of suffix marker occurrence according to the morpho-syntactic values by operating the longest suffix stripping methodology. The lemmatizer achieved an accuracy of 96.90% and experimented against a manually annotated test dataset.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Islam et al. (2024) studied this question.

synapsesocial.com/papers/68e6beabb6db64358763f1aahttps://doi.org/10.1109/iceeict62016.2024.10534443
Ask AI
Helpful
Bookmark
Share
View Full Paper