Synapse
⌘+K
Synapse
PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
June 20, 2026International Journal of Asian Language Processing

Leveraging LLM-generated Synthetic Data for Low-Resource Named Entity Recognition in Historical Japanese Documents

View Full Paper
Ask AI
Bookmark
Share

Authors

BWBohao WuAMAkira MaedaRARyo Akama

Discussion

Loading...

Member takes

Overview

Randomized trial investigates synthetic data improvement in named entity recognition for historical Japanese texts, suggesting effective methodologies.

Key Points

  • This work aims to enhance named entity recognition in historical Japanese documents by leveraging synthetic data generated by large language models.
  • Investigated NER on Yakusha Hyōbanki, a collection of historical texts.
  • Developed a multi-stage training framework including masked language modeling and NER training on synthetic data.
  • Conducted experiments across seven model architectures with varying synthetic data scales.
  • Synthetic-data augmentation consistently improved NER performance over a baseline method.
  • Limited gains were observed from domain-adaptive pretraining when high-quality synthetic data was used, highlighting a trade-off between cost and accuracy.
  • A two-stage synthetic-to-real pipeline was found to be an effective strategy for low-resource historical Japanese NER.

Cite This Study

Wu et al. (2026) studied this question.

synapsesocial.com/papers/6a3632d2db0793dc1a5394afhttps://doi.org/10.1142/s2717554526500086
View Full Paper
Ask AI
Bookmark
Share