Key points are not available for this paper at this time.
Due to the scarcity of annotated data, NLP tools, and pre-trained models, developing reliable natural language generation (NLG) systems for low-resource languages continues to be a major issue. This paper presents a comprehensive evaluation of four Transformer-based architectures for the news headline generation in the Nepali, a morphologically rich and underrepresented language. We evaluated four models representing different architectural paradigms: mBART-large-cc25 and mT5-small (multilingual encoder–decoder models), NepBERTa (a monolingual encoder model adapted to a sequence-to-sequence architecture), and LLaMA-3.2-1B (a multilingual decoder-only model, not finetuned in Nepali). All models were fine-tuned on a custom-curated dataset (NEPHEAD) under a unified experimental setup to enable controlled comparison. Model performance was evaluated using both lexical and semantic automatic metrics, that include ROUGE, BLEU, METEOR, BERTScore, and SBERT-based similarity, along with human evaluation based on relevance, fluency, conciseness, accuracy, and engagement. Inter-annotator agreement using Spearman’s rank correlation coefficient ( ρ ) and Pearson correlation coefficient (r) was also ccalculated to assess the reliability of evaluation. The results show that multilingual encoder–decoder models, particularly mBART-large-cc25, achieve the best overall performance in both automatic and human evaluation. The mT5-small model performs competitively, while LLaMA-3.2-1B demonstrates strong adaptability despite the absence of explicit Nepali pretraining. In contrast, NepBERTa exhibits limited effectiveness due to architectural and tokenization constraints. These findings highlight the importance of pretraining objectives, model architecture, and tokenizer design in low-resource text generation and provide practical insight to develop reliable NLG systems in underrepresented languages.
Dahal et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: