PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 23, 2026Procedia Computer Science0 citationsOpen Access

Saudi Dialects to MSA Machine Translation: A Systematic Evaluation of LLMs

View Full Paper
GAGhada Alharbi

Key Points

  • This research aims to evaluate the effectiveness of large language models in translating Saudi Dialectal Arabic to Modern Standard Arabic.
  • Two LLM models evaluated: ALLaM (multidialectal) and GPT-3.5 (multilingual)
  • Designed a specific prompt framework for model evaluation
  • Used two dialect datasets: SauDial and MADAR
  • Assessed performance with BLEU, TER, METEOR, and COMET metrics.
  • ALLaM outperforms GPT-3.5 in translating both datasets
  • Average BLEU score achieved: 53.17% for SauDial and 35.02% for MADAR
  • Demonstrated effectiveness of LLMs for low-resource Arabic dialect translation

Abstract

Converting Dialectal Arabic (DA) texts into Modern Standard Arabic (MSA) is considered an important step in downstream applications. This is because the huge diversity of Arabic dialects leads to limited available resources of such data. Besides, dealing directly with DA involves multiple complexities, as there are no specialized preprocessing tools designed for this form of the language. The recent advances in Large Language Models (LLMs) sound like a promising technology to ease the process of translating DA texts to MSA, particularly for such a low-resource task. In this paper, we evaluate the efficacy of LLMs on the task of translating fine-grained Saudi Dialectal Arabic to MSA. This was done using two LLM-based models, namely ALLaM (a multidialectal model) and GPT-3.5 (a multilingual model). This process involved designing a specific prompt framework for the LLM models. The experiments were conducted on two different Saudi dialect datasets: SauDial (a newly developed dataset) and MADAR (a benchmark dataset). The results were evaluated using four metrics: BLEU, TER, METEOR, and COMET. Our findings indicate that ALLaM significantly outperforms GPT-3.5 on both datasets across different Saudi dialects, with an average BLEU score of 53.17% and 35.02% on SauDial and MADAR, respectively.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ghada Alharbi (2026) studied this question.

synapsesocial.com/papers/69c0ddb8fddb9876e79c125dhttps://doi.org/10.1016/j.procs.2026.01.114
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Diglossia1959 · 2,983 citations
  2. 2AraT5-MSAizer: Translating Dialectal Arabic to MSA2024 · 2 citations
  3. 3SauDial: The Saudi Arabic dialects game localization dataset2025 · 3 citations
  4. 4Farasa: A Fast and Furious Segmenter for Arabic2016 · 375 citations
  5. 5Sirius_Translators at OSACT6 2024 Shared Task: Fin-tuning Ara-T5 Models for Translating Arabic Dialectal Text to Modern Standard Arabic2024 · 3 citations