In general, methods based on dictionaries performed better with PHI that is rarely mentioned in clinical text, but are more difficult to generalize. Methods based on machine learning tend to perform better, especially with PHI that is not mentioned in the dictionaries used. Finally, the issues of anonymization, sufficient performance, and "over-scrubbing" are discussed in this publication.
No takes yet. Share an insight, caveat, or question.
Meystre et al. (2010) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: