PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 2026Cancers4 citationsOpen Access

Revisiting AI Interpretability in Precision Oncology: Why Predictive Accuracy Does Not Ensure Stable Feature Importance

SOSouichi OkaYTYoshiyasu Takefuji

Key Points

  • This research aims to reassess AI interpretability in precision oncology by focusing on the stability of feature importance rankings.
  • Analyzed The Cancer Genome Atlas (TCGA) breast cancer multi-omics dataset.
  • Compared supervised models (Linear Regression, LASSO, Random Forest, XGBoost) with unsupervised and statistical methods.
  • Evaluated model feature rankings and stability upon minimal input perturbations.
  • Used Random Forest classifier with stratified 10-fold cross-validation for predictive performance.
  • Supervised models showed unstable feature importance rankings even with minimal perturbations (<0.1% feature removal).
  • High predictive accuracy often concealed fragile or misleading feature explanations.
  • Highly Variable Gene Selection and Spearman’s correlation produced stable feature sets and maintained competitive predictive performance.

Abstract

Background: Artificial intelligence (AI) is becoming important in oncology, supporting risk prediction, treatment planning, and biomarker discovery. However, current evaluation practices often assume that high predictive accuracy implies reliable interpretation—a misconception that may undermine reproducibility and clinical decision-making. This study aims to reassess interpretability by introducing feature ranking order consistency as a stability-focused metric to evaluate how model explanations respond to minimal input perturbations. Methods: Using The Cancer Genome Atlas (TCGA) breast cancer multi-omics dataset, we compared supervised models—Linear Regression, Least Absolute Shrinkage and Selection Operator (LASSO), Random Forest, and Extreme Gradient Boosting (XGBoost)—with unsupervised and statistical methods, including Principal Component Analysis (PCA), Highly Variable Gene Selection, and Spearman’s rank correlation. Each method produced a Top 20 feature ranking, and stability was assessed by testing whether rankings remained consistent after removing the top-ranked feature. Predictive performance was evaluated using a Random Forest classifier with stratified 10-fold cross-validation. Results: Supervised models exhibited unstable feature importance rankings even under minimal perturbations (<0.1% feature removal), suggesting that high predictive accuracy may obscure fragile or misleading explanations. In contrast, Highly Variable Gene Selection and Spearman’s correlation consistently produced stable, biologically coherent feature sets and maintained competitive predictive performance. Conclusions: Interpretive instability is a major limitation of many machine learning models in oncology. Incorporating stability-based criteria—such as feature ranking consistency—into evaluation frameworks is essential for ensuring reproducible, trustworthy, and clinically actionable AI. As AI adoption accelerates, prioritizing interpretability alongside accuracy is critical for responsible deployment in precision oncology.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Oka et al. (2026) studied this question.

synapsesocial.com/papers/698d6f0d5be6419ac0d551d4https://doi.org/10.3390/cancers18040593
Ask AI
Helpful
Bookmark
Share
View Full Paper