The ‘rojak’ language, a complex linguistic blend of Malay, English, Mandarin, and local dialects prevalent in Malaysia and Singapore, presents significant translation challenges due to its informal expressions, frequent code-switching, and cultural idioms. Although LLMs have shown growing capability in multilingual translation but there is limited works to evaluate the performance of LLMs under ‘rojak’ context. Furthermore, bilingual datasets are common but not ‘rojak’ dataset. This study tries to reduce the research gap by 1) proposing a ‘rojak’ dataset that capture informal slang, abbreviations and emoticons, and 2) providing early finding on how prompt engineering and fine-tuning strategies affect LLMs, particularly LLaMA model in translation from ‘rojak’ language to English. Evaluation was performed using metrics, including BLUE, BERTScore, METEOR, TER and COMET. the comparable performance of the fine-tuned LLaMA 3 8B highlights that parameter-efficient adaptation of open models can still yield competitive quality for low-resource, hybrid languages, demonstrating the potential of targeted fine-tuning in multilingual translation research.
Tan et al. (Thu,) studied this question.