At the expense of quantity and quality of training data, corpus-based models are becoming superior to rule-based models in solving complex Natural Language Processing (NLP) problems. In this research work, three categories of Machine Learning (ML) models for Machine Translatiteration (MTx) tasks are examined in a strictly low-resource scenario. This work studies and compares the Rule-Based Machine Transliteration (RBMTx) model, Statistical Machine Transliteration (SMTx) model and five generation-defining Neural Machine Transliteration (NMTx) models for the transliteration task in the low-resource Manipuri language. The work also discusses the contemporary script issues for the Manipuri language. The work explored and demonstrated how existing RBMTx models can facilitate corpus-based data-intensive machine learning models for low-resource languages using a novel technique for building parallel datasets. This study produced a gold-standard corpus of 35,000 Bengali script-Meetei Mayek parallel Manipuri words. With a Character Error Rate (CER) of only 0.66, a chrF score of 98.3, a BLEU score of 98.7 and a METEOR score of 99.00, the best performing Encoder-Decoder with Self Attention machine transliteration model sets a new performance record for the Bengali script to Meetei Mayek transliteration task. In addition to its immense potential for facilitating the ongoing script transition from Bengali script to Meetei Mayek, this research work will also help in addressing the low-resource bottleneck of Meetei Mayek for downstream Manipuri language NLP tasks.
Moirangthem et al. (2026) studied this question.