Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
July 14, 2026Open Access

LLM-Generated Synthetic Data for Low-Resource Urdu–English Neural Machine Translation

View Full Paper
Ask AI
Bookmark
Share

Authors

MUMuhammad Umer

Discussion

Loading...

Member takes

Overview

Randomized trial investigates synthetic data's effect on translation accuracy in low-resource Urdu–English context, suggesting improved model performance.

Key Points

  • To assess whether LLM-generated synthetic data can enhance neural machine translation for Urdu–English, given the lack of quality parallel corpora.
  • Utilized a fine-tuned mT5-small model on 19,793 Urdu–English sentence pairs from OPUS-100.
  • Generated synthetic data through back-translation of monolingual English sentences into Urdu and paraphrasing of existing English source sentences.
  • Trained both models under identical architectures to isolate effects of added data.
  • The augmented model improved BLEU scores from 7.37 to 8.21 and chrF scores from 25.02 to 26.10.
  • Error analysis revealed that improvements were primarily observed in longer source sentences.
  • Synthetic data enhanced generalization for challenging inputs rather than producing uniform gains.

Cite This Study

Muhammad Umer (2026) studied this question.

synapsesocial.com/papers/6a55d16e5aafca87247f860fhttps://doi.org/10.5281/zenodo.21326263
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Towards Multilingual Machine Translation for Low-Resource South Asian Languages: A Transformer-Based Approach on English–Urdu–Kashmiri2025
  2. 2Generalists vs. Specialists: Evaluating Large Language Models for Urdu2024
  3. 3Multilingual Generative AI Framework for Urdu and Regional Language Understanding Using Large Language Models2026
  4. 4Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks2024
  5. 5Improving Language Models Trained on Translated Data with Continual Pre-Training and Dictionary Learning Analysis2024