PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 15, 2005127 citations

Detecting phrase-level duplication on the world wide web

View Full Paper
DFDennis FetterlyGoogle (United States)MMMark S. ManasseSan Diego Mesa CollegeMNMarc NajorkGoogle (United States)

Key Points

Key points are not available for this paper at this time.

Abstract

Two years ago, we conducted a study on the evolution of web pages over time. In the course of that study, we discovered a large number of machine-generated "spam" web pages emanating from a handful of web servers in Germany. These spam web pages were dynamically assembled by stitching together grammatically well-formed German sentences drawn from a large collection of sentences. This discovery motivated us to develop techniques for finding other instances of such "slice and dice" generation of web pages, where pages are automatically generated by stitching together phrases drawn from a limited corpus. We applied these techniques to two data sets, a set of 151 million web pages collected in December 2002 and a set of 96 million web pages collected in June 2004. We found a number of other instances of large-scale phrase-level replication within the two data sets. This paper describes the algorithms we used to discover this type of replication, and highlights the results of our data mining.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fetterly et al. (2005) studied this question.

synapsesocial.com/papers/6a160781f239148d18c5b6e4https://doi.org/10.1145/1076034.1076066
Ask AI
Helpful
Bookmark
Share
View Full Paper