Randomized trial demonstrates high clinical sensitivity in lung cancer classification, suggesting a novel diagnostic tool.
Transcriptomic classification based on high-dimensional genomic datasets holds immense potential for precision diagnostics, yet statistical pitfalls like patient-level data leakage and technical batch effects have hindered the clinical translation of machine learning (ML) models. This study introduces CytoGraph-ML, an end-to-end, zero-leakage machine learning framework designed to classify cancer presence (Tumor vs. Normal tissue) while providing game-theoretic interpretability and direct mapping of predictive genomic features to documented biological pathways. Leveraging real-world clinical cohorts from the Gene Expression Omnibus (GEO), we trained our pipeline on a Lung Adenocarcinoma cohort (GSE10072, N=107) and validated it on an entirely independent external cohort (GSE19804, N=120) and a colorectal cohort (GSE21510, N=148). To ensure statistical integrity, we constructed a modular, leakage-free pipeline integrating Quantile Normalization, Group-Blind validation (via GroupShuffleSplit and GroupKFold on Patient IDs) to prevent patient-level leakage, a 14-gene biological proxy blacklist (removing VWF, PECAM1, CD34, IL6, IL8, and GAPDH to force the model to ignore stromal/inflammatory shortcuts), and non-parametric Mutual Information feature selection. We evaluated a deterministic Random Forest (RF) Classifier optimized for clinical sensitivity using a 5:1 class weight in favor of the Tumor class. The production model achieved an "In-Study" cross-validation accuracy of 98.57% (±2.86%) on GSE10072 and a final holdout test accuracy of 100%. Subjected to the cross-study "Acid Test" (GSE10072 to GSE19804) under platform-shift conditions, the safety-biased model achieved an accuracy of 50.83% but maintained a perfect clinical Sensitivity (Recall) of 1.0000 (zero false negatives) and an F1-score of 0.6704. Cross-tissue generalizability was further validated on colorectal cohort GSE21510, achieving 100% cross-validation and holdout accuracy, while label permutation audits yielded random-chance performance (42.42%), verifying the absence of leakage. Post-hoc explainability was established via TreeExplainer SHAP (Shapley Additive exPlanations) values, successfully mapping mathematical attributions to Reactome and KEGG pathways to isolate critical genomic drivers, including LDB2 (suppressor signaling), SLIT3 (mediating cell migration and angiogenesis), EPAS1 (adapting to hypoxia), EDNRB (G-protein coupled signaling), and MCL1 (anti-apoptotic survival). Finally, we address translational challenges, presenting a containerized FastAPI inference service (/predict) that implements Reference-Based Quantile Normalization (R-QN) to resolve the single-sample inference gap under batch effects. CytoGraph-ML demonstrates that transcriptomic classification can be executed with high clinical sensitivity while simultaneously unlocking the mathematical "black box" of machine learning, bridging the gap between computational accuracy and clinical interpretability. (FOR RESEARCH USE ONLY (RUO). NOT FOR USE IN DIAGNOSTIC PROCEDURES. MIT License.)
No takes yet. Share an insight, caveat, or question.
Mainak Biswas (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: