Synapse
⌘+K
Synapse
PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
August 13, 2026ACM Transactions on Asian and Low-Resource Language Information ProcessingOpen Access

Sequence Labeling in Urdu Social Media Texts: Data Annotation and Transformer-Based Deep Learning Models

View Full Paper
Ask AI
Bookmark
Share

Discussion

Loading...

Member takes

Overview

Randomized trial evaluates sequence labeling performance in Urdu texts, suggesting advancements for low-resource languages.

Key Points

  • This research aims to enhance sequence labeling tasks for Urdu social media texts by creating annotated datasets and benchmarking methods.
  • Introduced two annotated datasets for Urdu tweets: a POS tagging dataset with 39 tags and an NER dataset for three entity types.
  • Utilized a range of methods from traditional CRFs to deep learning techniques, including a transformer-based hybrid architecture.
  • Proposed XLM-R–CNN–BiLSTM–CRF for optimal sequence labeling and benchmarked its performance against various baselines.
  • Achieved an F1 score of 95.39% for POS tagging and 91.91% for NER using the proposed model.
  • Performance significantly surpassed strong baseline models in sequence labeling tasks.
  • Demonstrated the efficacy of using transformer-based architectures for low-resource language processing.

Cite This Study

A 2026 study studied this question.

synapsesocial.com/papers/6a7d75d82b0e0cff3f63ebc1https://doi.org/10.1145/3837066
View Full Paper
Ask AI
Bookmark
Share