This study introduces a novel method for de-identifying Portuguese clinical narratives by integrating transformer-based models with rule-based techniques. A BioBERTpt model was fine-tuned using a corpus of clinical cardiology and pulmonology texts. The model combining BioBERTpt and regular expressions achieved superior precision (0.92), recall (0.93), and F1-scores (0.93), significantly outperforming baseline models. The approach ensures data utility while complying with privacy regulations, highlighting its potential for clinical text anonymization in underrepresented languages.
Schneider et al. (Thu,) studied this question.