This paper presents a clustering technique that reduces the susceptibility to data noise by learning and clustering the data-distribution and then assigning the data to the cluster of its distribution. In the process, it reduces the impact of noise on clustering results. This method involves introducing a new distance among distributions, namely the expectation distance (denoted, ED), that goes beyond the state-of-art distribution distance of optimal mass transport, also called 2-Wasserstein (denoted, <tex-math notation="LaTeX">W₂</tex-math> ): The latter essentially depends only on the marginal distributions while the former also employs the information about the joint distributions, making it more powerful. Using the ED, the paper extends the classical <tex-math notation="LaTeX">K</tex-math> -means and <tex-math notation="LaTeX">K</tex-math> -medoids clustering to those over data-distributions (rather than raw-data) and further introduces <tex-math notation="LaTeX">K</tex-math> -medoids using <tex-math notation="LaTeX">W₂</tex-math> . The paper also presents the closed-form expressions of the <tex-math notation="LaTeX">W₂</tex-math> and ED distance measures. The implementation results of the proposed ED and the <tex-math notation="LaTeX">W₂</tex-math> distance measures to cluster real-world weather data as well as stock data are also presented, which involves efficiently extracting and using the underlying data distributions—Gaussians for weather data versus lognormals for stock data. The results show striking performance improvement over classical clustering of raw-data, with higher accuracy realized for ED. Also, not only does the distribution-based clustering offer higher accuracy, but it also lowers the computation time due to reduced time-complexity.
No takes yet. Share an insight, caveat, or question.
Adesunkanmi et al. (2024) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: