PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 5, 2026International Journal of Advanced Computer Science and Applications0 citationsOpen Access

D-LexeCan: A Dynamic Lexicon-Based Framework for Sentiment Analysis in Tarifit, a Low-Resource Multiscript Language

AAAmar AmakssoumAbdelmalek Essaâdi UniversityFBFadwa BouhaferAbdelmalek Essaâdi UniversityAHAnass El HaddadiAbdelmalek Essaâdi University

Key Points

  • The aim is to develop an effective sentiment analysis framework for the Tarifit language, addressing challenges of low-resource data.
  • Introduced D-LexeCan framework leveraging dynamic lexicon induction.
  • Performed multi-script normalization unifying Arabic script, Tifinagh, and Arabizi.
  • Automatically induced sentiment-bearing unigrams and bigrams.
  • Modelled negation and amplification using linguistic operators.
  • Evaluated on a manually annotated social media corpus.
  • Achieved an accuracy of 0.8800 and a Macro-F1 score of 0.8798.
  • Outperformed static lexicon baseline (0.5275 accuracy) and classical models.
  • Demonstrated better performance than other neural architectures like BiLSTM (0.7950 accuracy).
  • Fine-tuned multilingual transformers (mBERT) reached an accuracy of 0.8175.

Abstract

Sentiment analysis for low-resource languages remains challenging due to limited annotated data, orthographic instability, informal writing practices, and the lack of dedicated linguistic resources, challenges that are particularly acute for Tarifit (Tamazight of the Rif), an under-resourced Amazigh language characterized by strong dialectal variation, pervasive multi-script usage, and highly noisy user-generated content on social media. This study introduces D-LexeCan, a dynamic lexicon-based sentiment analysis framework that infers polarity directly from annotated corpus evidence without relying on predefined sentiment dictionaries or computationally intensive pretrained deep learning and transformer-based models. The framework combines deterministic multi-script normalization, unifying Arabic script, Tifinagh, and Arabizi into a single Tarifit Latin representation with automatic induction of sentiment-bearing unigrams and bigrams, while explicitly modeling negation and amplification phenomena through linguistically motivated operators and preserving emojis as meaningful discourse-level sentiment cues. The approach is evaluated on a manually annotated social media corpus collected from multiple online platforms, where it achieves an accuracy of 0.8800 and a Macro-F1 score of 0.8798. The results outperform a static lexicon baseline with an accuracy of 0.5275, a classical machine-learning model based on TF–IDF and SVM with an accuracy of 0.8525, and neural architectures including BiLSTM with an accuracy of 0.7950. Experiments with frozen multilingual transformer encoders show accuracy ranging from 0.6725 to 0.7650. Fine-tuned multilingual transformers such as mBERT achieve competitive performance, reaching an accuracy of 0.8175. Overall, the results demonstrate that adaptive and linguistically grounded dynamic lexicon induction constitutes an effective, interpretable, and computationally efficient alternative for sentiment analysis in low-resource, noisy, and multi-script African language contexts.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Amakssoum et al. (2026) studied this question.

synapsesocial.com/papers/69d1fdd4a79560c99a0a4281https://doi.org/10.14569/ijacsa.2026.0170372
Ask AI
Helpful
Bookmark
Share
View Full Paper