Traditional statistical methods, such as principal component analysis, often fail to capture complex dependencies in gene expression data. To address this limitation, we propose a functional framework combining multidimensional scaling with a density-based version of principal component analysis. By representing gene expression profiles through estimated distributions, the method captures both distributional shape and variability across individuals. Using artificial datasets, we show that clustering performed on density-based scores accurately recovers the original class structure. We also compare our approach with two nonlinear dimensionality reduction techniques, uniform manifold approximation and projection and diffusion maps, as dimensionality increases, highlighting the importance of L2 normalization in preserving discriminative power. For moderate dimensions, densities are estimated using a multivariate gamma kernel, well suited to the non-negative and asymmetric nature of transcriptomic data. Finally, we establish convergence results for the estimated inner products and prove the spectral consistency of the resulting eigenvalues and eigenvectors. The extracted principal components effectively capture both lower-order statistical moments and complex gene interaction patterns that are often inaccessible to classical linear methods.
Smail Yousfi (Wed,) studied this question.