PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 3, 2026IEEE Access0 citationsOpen Access

A Survey of Text Deduplication: From Syntactic Matching to Semantic Understanding

View Full Paper
JKJeongmin KimKangwon National UniversityYPYeonsu ParkEwha Womans University

Key Points

  • Identifying and eliminating redundant text improves data quality, optimizing resources and AI model performance.
  • The survey covers both foundational and modern techniques, including shingling, LSH, and embeddings with Transformers.
  • A comprehensive analysis includes applications, evaluation methodologies, and benchmarks in text deduplication.
  • Challenges persist in scaling systems for petabyte-level data, with open questions around semantic intelligence and fairness.

Abstract

The exponential growth of digital text has made text deduplication, the process of identifying and eliminating redundant information, a critical task for enhancing data quality, optimizing computational resources, and improving the performance of downstream applications, particularly the training of massive AI models. This survey provides a comprehensive and structured overview of the field, beginning with a formal taxonomy of duplicate types including exact, near, and semantic, and a detailed breakdown of the core deduplication pipeline from preprocessing to similarity metrics. We then systematically review the evolution of algorithmic approaches, covering foundational syntactic techniques such as shingling and Locality Sensitive Hashing (LSH), as well as modern semantic methods driven by text embeddings, Transformers, and Large Language Models (LLMs). Beyond algorithms, the survey addresses the critical aspects of system design and scalability for data at the petabyte scale, including architectural patterns, distributed processing, and considerations for dynamic, streaming environments. A detailed examination of diverse applications, rigorous evaluation methodologies, and standard benchmarks provides a practical context for these techniques. Finally, we synthesize the persistent challenges and identify key open research questions, culminating in a visionary outlook on future research directions centered on advanced semantic intelligence, deduplication across multiple modalities and languages, and the engineering of trustworthy, private, and fair systems. This work serves as an essential reference for both newcomers and experienced researchers, providing a complete roadmap of the text deduplication landscape from foundational principles to the frontiers of analysis driven by Artificial Intelligence (AI).

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kim et al. (2026) studied this question.

synapsesocial.com/papers/69a75d19c6e9836116a26914https://doi.org/10.1109/access.2026.3658439
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1From Foundations to GPT in Text Classification: A Comprehensive Survey on Current Approaches and Future Trends2025 · 17 citations
  2. 2Improving Academic Plagiarism Detection for STEM Documents by Analyzing Mathematical Content and Citations2019 · 24 citations
  3. 3Long Short-Term Memory1997 · 101,538 citations
  4. 4A comprehensive survey of text classification techniques and their research applications: Observational and experimental insights2024 · 68 citations
  5. 5Large language models: an overview of foundational architectures, recent trends, and a new taxonomy2025 · 37 citations