PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 8, 2007617 citations

Detecting near-duplicates for web crawling

View Full Paper
GMGurmeet Singh MankuAJArvind Kumar JainASAnish Das Sarma

Key Points

Key points are not available for this paper at this time.

Abstract

Near-duplicate web documents are abundant. Two such documents differ from each other in a very small portion that displays advertisements, for example. Such differences are irrelevant for web search. So the quality of a web crawler increases if it can assess whether a newly crawled web page is a near-duplicate of a previously crawled web page or not. In the course of developing a near-duplicate detection system for a multi-billion page repository, we make two research contributions. First, we demonstrate that Charikar's fingerprinting technique is appropriate for this goal. Second, we present an algorithmic technique for identifying existing f-bit fingerprints that differ from a given fingerprint in at most k bit-positions, for small k. Our technique is useful for both online queries (single fingerprints) and all batch queries (multiple fingerprints). Experimental evaluation over real data confirms the practicality of our design.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Manku et al. (2007) studied this question.

synapsesocial.com/papers/6a0ffd3c2badbc352aff18e4https://doi.org/10.1145/1242572.1242592
Ask AI
Helpful
Bookmark
Share
View Full Paper