PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 2026Scientific Reports1 citationsOpen Access

Deep learning enabled pseudonymization for preserving data privacy of financial identifiers in public documents in India

View Full Paper
RRR. RoopalakshmiManipal Academy of Higher EducationSKSaurabh KailasManipal Academy of Higher EducationRSR. Sreelatha

Key Points

  • The study aims to develop a deep learning framework for pseudonymizing handwritten signatures in public documents to enhance data privacy.
  • Developed a Fully Convolutional Neural Network (CNN)-based framework using SuperPoint architecture.
  • Integrated Differentiable output decoding for optimal pseudonymization.
  • Evaluated the framework on a dataset of over 500 real-world Permanent Account Number (PAN) cards.
  • Compared performance with traditional techniques such as ORB, FAST, SIFT, and deep CNN.
  • Demonstrated superior performance in precision and recall metrics compared to baseline techniques.
  • Achieved efficient runtime and minimal spatial overhead for processing documents.
  • Highlighted practical implications for secure digital archiving in compliance with GDPR standards.

Abstract

Abstract The increasing digitization and transmission of government-issued electronic documents have intensified the need to protect the ’Handwritten signatures’-recognized as ’critical biometric identifiers’ from identity-related data breaches. For instance, as per 2025-RSA ID IQ Report, 40% of respondents reported Identity-related data breaches and 66% emphasized the significant damages caused by these breaches to their organizations. The existing privacy-preserving anonymization research is primarily focusing on facial features and fingerprints, whereas the Pseudonymization of handwritten signatures in publicly accessible documents remains largely underexplored in the literature. This research study proposes a new Fully Convolutional Neural Network (CNN)-based Pseudonymization framework using SuperPoint architecture integrated with Differentiable output decoding, which aims to identify and pseudonymize the handwritten signatures in public-domain documents, specifically in Indian Government issued Permanent Account Number (PAN) cards. In contrast to traditional anonymization approaches, this pseudonymization technique preserves document utility by securing sensitive data and thereby enables traceable identity protection without compromising the structural integrity of input documents. Extensive evaluations are carried out on a curated dataset of 500+ real-world PAN cards, which establishes the model’s robustness and applicability in large-scale deployments. The results of comparative analysis with baseline techniques including ORB, FAST, SIFT and deep CNN, clearly demonstrate the superior performance of the proposed method in terms of various metrics such as Precision, Recall, SSIM, runtime efficiency, and spatial overhead. In addition, the research findings suggest practical implications for embedding CNN-based pseudonymization into Public-sector Document processing pipelines, which supports secure and utility-preserved digital archiving in compliance with modern privacy GDPR standards.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Roopalakshmi et al. (2026) studied this question.

synapsesocial.com/papers/698d6e6e5be6419ac0d541achttps://doi.org/10.1038/s41598-026-39309-6
Ask AI
Helpful
Bookmark
Share
View Full Paper