Randomized trial investigates synthetic data improvement in named entity recognition for historical Japanese texts, suggesting effective methodologies.
Key Points
This work aims to enhance named entity recognition in historical Japanese documents by leveraging synthetic data generated by large language models.
Investigated NER on Yakusha Hyōbanki, a collection of historical texts.
Developed a multi-stage training framework including masked language modeling and NER training on synthetic data.
Conducted experiments across seven model architectures with varying synthetic data scales.
Synthetic-data augmentation consistently improved NER performance over a baseline method.
Limited gains were observed from domain-adaptive pretraining when high-quality synthetic data was used, highlighting a trade-off between cost and accuracy.
A two-stage synthetic-to-real pipeline was found to be an effective strategy for low-resource historical Japanese NER.