PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 5, 2026Scientific Reports3 citationsOpen Access

ModernBERT is more efficient than conventional BERT for chest CT findings classification in Japanese radiology reports

YYYosuke YamagishiTKTomohiro KikuchiSHShouhei Hanaoka

Key Points

  • This research aims to evaluate the efficiency and performance of ModernBERT compared to traditional BERT models in classifying chest CT findings within Japanese radiology reports.
  • Compared three Japanese models: BERT Base, JMedRoBERTa, ModernBERT.
  • Fine-tuned all models on the CT-RATE-JPN dataset under identical conditions.
  • Constructed an external dataset, RR-Findings, to test generalizability.
  • ModernBERT required fewer tokens and resulted in faster training and inference.
  • Exact match accuracy was 74.7% for ModernBERT vs. 72.7% for BERT Base on the internal dataset.
  • In the external dataset, BERT Base surpassed both JMedRoBERTa and ModernBERT, with ModernBERT showing the largest decline in performance.

Abstract

Japanese language models for medical text classification face challenges with complex vocabulary and linguistic structures in radiology reports. This study compared three Japanese models—BERT Base, JMedRoBERTa, and ModernBERT—for multi-label classification of 18 chest CT findings. Using the CT-RATE-JPN dataset, all models were fine-tuned under identical conditions. ModernBERT showed clear efficiency advantages, producing substantially fewer tokens and achieving faster training and inference than the other models while maintaining comparable performance on the internal test dataset (exact match accuracy: 74.7% vs. 72.7% for BERT Base). To assess generalizability, we additionally constructed RR-Findings, an external dataset of 243 naturally written Japanese radiology reports annotated using the same schema. Under this domain-shifted setting, performance differences became pronounced: BERT Base outperformed both JMedRoBERTa and ModernBERT, whereas ModernBERT showed the largest decline in exact match accuracy. Average precision differences were smaller, indicating that ModernBERT retained reasonable ranking ability despite reduced calibration. Overall, ModernBERT offers substantial computational efficiency and strong in-domain performance but remains sensitive to real-world linguistic variability. These results highlight the need for more diverse natural-language training data and domain-specific calibration strategies to improve robustness when deploying modern transformer models in heterogeneous clinical environments.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yamagishi et al. (2026) studied this question.

synapsesocial.com/papers/69d1fd62a79560c99a0a36c7https://doi.org/10.1038/s41598-026-44292-z
Ask AI
Helpful
Bookmark
Share
View Full Paper