Type 2 diabetes (T2D), inflammatory bowel disease (IBD), and colorectal cancer (CRC) share overlapping metabolic alterations that hinder early, disease-specific diagnosis. Using publicly available serum metabolomics data sets (T2D: ST003390; IBD: ST003312; CRC: ST000284), a standardized workflow combining random forest-based imputation, log transformation, Pareto scaling, and ComBat batch correction was implemented prior to supervised machine learning. Eight algorithms (logistic regression, linear and RBF SVM, random forest, XGBoost, k-nearest neighbors, multilayer perceptron, and partial least-squares-discriminant analysis) were benchmarked for binary and multiclass classification using stratified 5-fold cross-validation, F1-scores, and bootstrapped ROC-AUC estimates. Binary models yielded near-perfect discrimination for T2D (AUC ≈ 1.0) and high accuracy for IBD and CRC (AUC 0.93-0.95), while multilayer perceptron and partial least-squares-discriminant analysis achieved multiclass accuracy >0.9 and macro-AUC 0.98. Mapping discriminative metabolites to KEGG pathways revealed disease-linked signatures, including glucose and lipid metabolism in T2D, amino acid and porphyrin metabolism in IBD, and nucleotide and sphingolipid metabolism in CRC, supporting proteome-metabolome network perturbations. The current comparative machine learning framework of serum metabolome demonstrates a robust, though variable, multidisease classification performance across conditions (T2D, IBD, and CRC) used in this study. This strategy has the potential to provide interpretable pathway-level markers that may inform future proteome- and metabolome-centered diagnostic strategies.
Das et al. (Thu,) studied this question.