Synapse
⌘+K
Synapse
PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
August 13, 2026ACM Transactions on Asian and Low-Resource Language Information ProcessingOpen Access

Qalb: Largest State-of-the-Art Urdu Large Language Model for 230M Speakers with Systematic Continued Pre-Training

View Full Paper
Ask AI
Bookmark
Share

Authors

MHMuhammad Taimoor HassanAuburn UniversityJAJawad AhmedInternational Islamic University, IslamabadMAMuhammad AwaisUniversity of Engineering and Technology Lahore

Discussion

Loading...

Member takes

Overview

Randomized trial shows improved Urdu text generation in a language model, suggesting better NLP performance for low-resource languages.

Key Points

  • The aim is to develop a state-of-the-art Urdu language model that effectively addresses the challenges in generating and understanding Urdu text.
  • Developed Qalb through continued pre-training on 1.97 billion token dataset, followed by supervised fine-tuning on Alif Urdu-instruct dataset.
  • Included diverse Urdu texts and 140 million tokens of English data to prevent catastrophic forgetting.
  • Evaluated the model across seven diverse tasks, including Classification and Sentiment Analysis.
  • Qalb achieved a weighted average score of 90.34 on Urdu benchmarks, outperforming Alif-1.0-Instruct by 3.24 points.
  • Surpassed the base LLaMA 3.1 8B-Instruct model by 44.64 points in Urdu text generation performance.
  • Demonstrated effective adaptation of foundation models to low-resource languages through comprehensive evaluations.

Cite This Study

Hassan et al. (2026) studied this question.

synapsesocial.com/papers/6a7d76652b0e0cff3f63faa1https://doi.org/10.1145/3839365
View Full Paper
Ask AI
Bookmark
Share