Key result
Five statistical classification methods (PLS, PPLS, LASSO, nearest shrunken centroids, and random forest) performed similarly in discriminating heart failure etiologies using gene expression profiles.
Why the study?
Can statistical methods accurately discriminate between ischemic and non-ischemic heart failure etiologies using gene expression profiles?
Can statistical methods accurately discriminate between ischemic and non-ischemic heart failure etiologies using gene expression profiles?
Several statistical methods perform similarly in discriminating heart failure etiology from gene expression data, highlighting the need for multiple or larger datasets for reliable classification.
Similar performance across methods cautions against clinical use for HF etiology discrimination; leaves open validation in larger gene expression datasets.
BACKGROUND: Human heart failure is a complex disease that manifests from multiple genetic and environmental factors. Although ischemic and non-ischemic heart disease present clinically with many similar decreases in ventricular function, emerging work suggests that they are distinct diseases with different responses to therapy. The ability to distinguish between ischemic and non-ischemic heart failure may be essential to guide appropriate therapy and determine prognosis for successful treatment. In this paper we consider discriminating the etiologies of heart failure using gene expression libraries from two separate institutions. RESULTS: We apply five new statistical methods, including partial least squares, penalized partial least squares, LASSO, nearest shrunken centroids and random forest, to two real datasets and compare their performance for multiclass classification. It is found that the five statistical methods perform similarly on each of the two datasets: it is difficult to correctly distinguish the etiologies of heart failure in one dataset whereas it is easy for the other one. In a simulation study, it is confirmed that the five methods tend to have close performance, though the random forest seems to have a slight edge. CONCLUSIONS: For some gene expression data, several recently developed discriminant methods may perform similarly. More importantly, one must remain cautious when assessing the discriminating performance using gene expression profiles based on a small dataset; our analysis suggests the importance of utilizing multiple or larger datasets.
No takes yet. Share an insight, caveat, or question.
Huang et al. (2005) studied Heart failure (n=66). Statistical classification methods (PLS, PPLS, LASSO, SC, RF) vs. Comparison between methods was evaluated on Leave-one-out cross-validation (LOOCV) misclassification errors. Five statistical classification methods (PLS, PPLS, LASSO, nearest shrunken centroids, and random forest) performed similarly in discriminating heart failure etiologies using gene expression profiles.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: