Background: Removing explicit protected health information (PHI) does not necessarily make a clinical narrative non-identifiable. Rare diagnoses, social circumstances, locations, temporal patterns, and treatment trajectories may still permit re-identification. Methods: This PRISMA-ScR-informed scoping review combined PubMed/MEDLINE, OpenAlex, and targeted NLP searches with OpenAlex title-and-abstract, Web of Science, and Scopus sensitivity checks. Eligible publications addressed clinical or biomedical text de-identification from 1 January 2018 to 31 March 2026. Results: The searches yielded 275 records and 258 unique records after deduplication. Eligibility assessment produced 110 core records, 11 background or review records, and 22 exclusions. All 143 eligibility-stage records underwent independent double screening and reconciliation; structured extraction covered all 110 core records. Abstract screening of 47 additional Scopus candidates found no task function or implementation paradigm outside the proposed classification. The resulting two-axis framework separates a method’s role in the workflow from its technical implementation. Conclusions: Clinical text de-identification cannot be reduced to named-entity recognition. Local hybrid or transformer pipelines remain the most defensible baseline for routine PHI detection. LLMs are better suited to defined tasks in augmentation, transformation, and assurance, with local validation and explicit control of data exposure.
Murat Sariyar (Fri,) studied this question.