PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 5, 2026PeerJ Computer Science0 citationsOpen Access

Multi-task fine-tuning using prefix-prepend format in text-to-text transfer transformer for automatic question generation in Bahasa Indonesia

View Full Paper
HAHalim Wildan AwalurahmanIBIndra Budi

Key Points

  • This research aims to improve automatic question generation in Bahasa Indonesia by addressing limitations of existing models.
  • Implemented prefix-prepend fine-tuning with the idT5-base model,
  • Utilized the Indonesian SQuAD and TyDiQA datasets for evaluation,
  • Evaluated model performance using BLEU, ROUGE, and BERT similarity metrics.
  • Achieved 0.1643 BLEU, 0.4099 ROUGE-L, 0.7177 BERT score on SQuAD,
  • Achieved 0.1941 BLEU, 0.4301 ROUGE-L, 0.7291 BERT score on TyDiQA,
  • Human evaluation indicated better performance than the baseline model.

Abstract

Automatic question generation is one solution to help create test items that require a lot of time and effort. The state-of-the-art model for automatic question generation in Bahasa Indonesia, which uses idT5, has several drawbacks, including misuse of context, overuse of question words, and an answer target that must be exactly from the context, which makes it extractive. This research aims to address those problems by improving the performance of previous models through a new fine-tuning scheme. This research proposed a prefix-prepend format for fine-tuning with the idT5-base model on the Indonesian Stanford Question Answering Dataset (SQuAD) and the Typologically Diverse Question Answering (TyDiQA) dataset. We evaluate the model with Bilingual Evaluation Understudy (BLEU), Recall-Oriented Understudy for Gisting Evaluation (ROUGE), and Bidirectional Encoder Representations from Transformers (BERT) similarity metrics. The results show that prefix-prepend fine-tuning improved the performance of the baseline model. Our best model achieved 0.1643 BLEU, 0.4099 ROUGE-L, and 0.7177 BERT similarity score on SQuAD, and 0.1941 BLEU, 0.4301 ROUGE-L, and 0.7291 BERT similarity score on TyDiQA. The human evaluation using the Content Validation Index (CVI) and a paired t-test indicated that the proposed model performed better than the baseline. While the proposed model addresses many of the baseline’s shortcomings, it still struggles to handle questions that require understanding complex relationships between entities. Future studies can explore improvements for this case using external knowledge or other models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Awalurahman et al. (2026) studied this question.

synapsesocial.com/papers/6a49f754f5d1d45b288013d3https://doi.org/10.7717/peerj-cs.3987
Ask AI
Helpful
Bookmark
Share
View Full Paper