Given a large collection of sparse vector data in a high dimensional space, we investigate the problem of finding all pairs of vectors whose similarity score (as determined by a function such as cosine distance) is above a given threshold. We propose a simple algorithm based on novel indexing and optimization strategies that solves this problem without relying on approximation methods or extensive parameter tuning. We show the approach efficiently handles a variety of datasets across a wide setting of similarity thresholds, with large speedups over previous state-of-the-art approaches.
No takes yet. Share an insight, caveat, or question.
Bayardo et al. (2007) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: