Benchmarking Dimensionality Reduction Methods for Livestock Transcriptomic Data: A Comparative Analysis of Visualization, Clustering, and Classification Performance
Comparative study demonstrates variable clustering and classification efficacy across reduction algorithms in livestock transcriptomics, highlighting t-SNE and PCA as top performers.
Key Points
To evaluate and benchmark the clustering, visualization, and classification performances of PCA, Kernel PCA, t-SNE, and UMAP on high-dimensional livestock transcriptomic datasets.
Evaluated four dimensionality reduction algorithms (PCA, Kernel PCA, t-SNE, and UMAP) across two independent livestock microarray datasets (GSE20552 and GSE24560) from the GEO database.
Assessed clustering performance using the Silhouette score, Davies–Bouldin index (DBI), and Calinski–Harabasz index (CHI).
Measured downstream classification performance using accuracy, area under the receiver operating characteristic curve (AUC), and F1-score.
On GSE20552, t-SNE achieved top clustering performance, while PCA yielded the highest classification metrics with an accuracy of 92.5%, AUC of 0.9750, and F1-score of 0.9278.
On GSE24560, t-SNE demonstrated superior performance in both clustering and classification tasks, obtaining an accuracy of 80.59%, AUC of 0.9165, and F1-score of 0.8057.
Overall algorithm performance ranked in descending order as t-SNE, PCA, Kernel PCA, and UMAP based on average rank values across evaluation metrics.