PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 10, 2026Engineering Reports0 citationsOpen Access

Towards Enhancing Healthcare Data Privacy: Integrating BioClinicalBERT With Polyalphabetic Cipher for Entity Recognition and Anonymization

View Full Paper
DSDeblina Mazumder SetuTITania IslamMRMM Rahman

Key Points

  • The aim is to develop a model that accurately identifies and anonymizes sensitive healthcare data while retaining its usefulness.
  • Preprocessing data with SpaCy for tokenization, lemmatization, and pattern recognition.
  • Integrating BioClinicalBERT and SpaCy for entity recognition using subword-level tokenization.
  • Optimizing model performance through hyperparameter tuning of learning rate, batch size, and epochs.
  • Implementing dual-layered encryption-based anonymization for comprehensive privacy protection.
  • Achieved a macro-average F1 score of 0.82 and a micro-average F1 score of 0.94.
  • Demonstrated high precision in recognizing sensitive entities like ID, religion, and marital status.
  • Effectively anonymized unstructured data and improved detection of complex compound entities.
  • Outperformed previous methods in preserving data utility for research while complying with privacy regulations.

Abstract

ABSTRACT Protecting sensitive data is important in healthcare to ensure privacy and follow rules like the General Data Protection Regulation (GDPR) and Health Insurance Portability and Accountability Act (HIPAA). However, existing studies struggle to handle unstructured data, miss sensitive information like compound entities, and reduce the usefulness of data during the anonymization process. This study aims to create a better model that can identify and anonymize sensitive data accurately and handle complex data. This work proposes an innovative method for preserving sensitive data in healthcare, incorporating natural language processing (NLP)‐based entity recognition and dual‐layered encryption‐based anonymization. The data is preprocessed with SpaCy for tokenization, lemmatization, and regex‐based pattern recognition to preserve formats like dates and IDs while identifying entities such as dates, diseases, and locations. Biomedical Clinical Bidirectional Encoder Representations from Transformers (BioClinicalBERT) and SpaCy are fine‐tuned with healthcare data, using subword‐level tokenization. Hyperparameter tuning optimizes performance by adjusting learning rate, batch size, and epochs. BioClinicalBERT, SpaCy, SciSpaCy, and regex patterns are then combined to detect sensitive information, which is anonymized using dynamic masking and encryption for comprehensive privacy protection. Our model performed well, with a macro‐average F1 score of 0.82, a micro‐average F1 score of 0.94, and high precision in recognizing sensitive entities like ID, religion, and marital status. This research also anonymizes unstructured data effectively. This outperforms existing methods by accurately detecting complex compound entities often missed by rule‐based systems. Moreover, compared to previous anonymization method, our dual‐layered anonymization process preserves data utility for research purposes. This research enhances patient data privacy in healthcare environments, ensuring compliance with GDPR and HIPAA regulations. It enables safer, more effective utilization of medical data for clinical research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Setu et al. (2026) studied this question.

synapsesocial.com/papers/69af952b70916d39fea4c74chttps://doi.org/10.1002/eng2.70673
Ask AI
Helpful
Bookmark
Share
View Full Paper