PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 19, 200998 citations

Brute force and indexed approaches to pairwise document similarity comparisons with MapReduce

View Full Paper
JLJimmy Lin

Key Points

Key points are not available for this paper at this time.

Abstract

This paper explores the problem of computing pairwise similarity on document collections, focusing on the application of "more like this" queries in the life sciences domain. Three MapReduce algorithms are introduced: one based on brute force, a second where the problem is treated as large-scale ad hoc retrieval, and a third based on the Cartesian product of postings lists. Each algorithm supports one or more approximations that trade effectiveness for efficiency, the characteristics of which are studied experimentally. Results show that the brute force algorithm is the most efficient of the three when exact similarity is desired. However, the other two algorithms support approximations that yield large efficiency gains without significant loss of effectiveness.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jimmy Lin (2009) studied this question.

synapsesocial.com/papers/6a1c221c0a1f7575939d9125https://doi.org/10.1145/1571941.1571970
Ask AI
Helpful
Bookmark
Share
View Full Paper