PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 1, 2007IEEE Transactions on Knowledge and Data Engineering317 citations

The Concentration of Fractional Distances

View Full Paper
DFDamien FrançoisVWVincent WertzMVMichel Verleysen

Key Points

Key points are not available for this paper at this time.

Abstract

Nearest neighbor search and many other numerical data analysis tools most often rely on the use of the euclidean distance. When data are high dimensional, however, the euclidean distances seem to concentrate; all distances between pairs of data elements seem to be very similar. Therefore, the relevance of the euclidean distance has been questioned in the past, and fractional norms (Minkowski-like norms with an exponent less than one) were introduced to fight the concentration phenomenon. This paper justifies the use of alternative distances to fight concentration by showing that the concentration is indeed an intrinsic property of the distances and not an artifact from a finite sample. Furthermore, an estimation of the concentration as a function of the exponent of the distance and of the distribution of the data is given. It leads to the conclusion that, contrary to what is generally admitted, fractional norms are not always less concentrated than the euclidean norm; a counterexample is given to prove this claim. Theoretical arguments are presented, which show that the concentration phenomenon can appear for real data that do not match the hypotheses of the theorems, in particular, the assumption of independent and identically distributed variables. Finally, some insights about how to choose an optimal metric are given.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

François et al. (2007) studied this question.

synapsesocial.com/papers/6a207446389e15e238f488b3https://doi.org/10.1109/tkde.2007.1037
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Introduction to Statistical Pattern Recognition1990 · 11,139 citations
  2. 2Relevance feedback: a power tool for interactive content-based image retrieval1998 · 1,774 citations
  3. 3Laws of Large Numbers for Pairwise Independent Uniformly Integrable Random Variables1987 · 26 citations
  4. 4The hybrid tree: an index structure for high dimensional feature spaces1999 · 207 citations
  5. 5Density-based indexing for approximate nearest-neighbor queries1999 · 96 citations