PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 23, 2026Procedia Computer Science0 citationsOpen Access

Towards a Translation Framework for ‘rojak’ Language: Challenges and Early Findings

View Full Paper
JTJia TanPGPey Yun GohSTShing Chiang Tan

Key Points

  • The study aims to address translation challenges in rojak language and evaluate LLM performance in this context.
  • Proposed a rojak dataset with informal slang, abbreviations, and emoticons.
  • Explored prompt engineering and fine-tuning strategies for the LLaMA model.
  • Evaluated translation performance using BLUE, BERTScore, METEOR, TER, and COMET metrics.
  • The fine-tuned LLaMA 3 8B model demonstrated competitive quality in translating low-resource languages.
  • Targeted fine-tuning showed promise for improving multilingual translation outcomes.

Abstract

The ‘rojak’ language, a complex linguistic blend of Malay, English, Mandarin, and local dialects prevalent in Malaysia and Singapore, presents significant translation challenges due to its informal expressions, frequent code-switching, and cultural idioms. Although LLMs have shown growing capability in multilingual translation but there is limited works to evaluate the performance of LLMs under ‘rojak’ context. Furthermore, bilingual datasets are common but not ‘rojak’ dataset. This study tries to reduce the research gap by 1) proposing a ‘rojak’ dataset that capture informal slang, abbreviations and emoticons, and 2) providing early finding on how prompt engineering and fine-tuning strategies affect LLMs, particularly LLaMA model in translation from ‘rojak’ language to English. Evaluation was performed using metrics, including BLUE, BERTScore, METEOR, TER and COMET. the comparable performance of the fine-tuned LLaMA 3 8B highlights that parameter-efficient adaptation of open models can still yield competitive quality for low-resource, hybrid languages, demonstrating the potential of targeted fine-tuning in multilingual translation research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tan et al. (2026) studied this question.

synapsesocial.com/papers/69c0de74fddb9876e79c1368https://doi.org/10.1016/j.procs.2026.01.046
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond2024 · 490 citations
  2. 2Large language models (LLMs): survey, technical frameworks, and future challenges2024 · 360 citations
  3. 3ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs2024 · 9 citations
  4. 4Prompting Multilingual Large Language Models to Generate Code-Mixed Texts: The Case of South East Asian Languages2023 · 27 citations
  5. 5Mandarin–English code-switching speech corpus in South-East Asia: SEAME2015 · 67 citations