This paper presents our participation in the NTCIR-18 RadNLP 2024 English main task and subtask. We describe our proposed solution to address the problem and discuss the official results. Our approach is based on large language models, with additional experiments involving data augmentation, retrieval-augmented generation, and prompting for the main task. Additionally, for the subtask, we employed a ModernBERT model with pre-training and hyperparameter optimization. Our best-performing submission in the main task, scores 0.5309\% in overall joint accuracy (fine) evaluation. Also, our best-performing submission in the subtask, scores 0.8189\% in overall micro F2.0 evaluation. Results from additional runs also show that data augmentation could further improve model performance beyond our best submission.
Díaz-Galiano et al. (2025) studied this question.