PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 9, 2026Interdisciplinary Journal of Information Knowledge and Management13 citationsOpen Access

Federated Tree-Based Ensembles with SHAP Explainability and Integrated Feature Selection for Secure Lung Cancer Health Analytics

VSVahiduddin ShariffCPChiranjeevi ParitalaKAKrishna Mohan Ankala

Key Points

  • The aim is to develop a privacy-preserving, explainable framework for health risk prediction using federated learning.
  • Integrated TreeNetX with SHAP-based explainability for model training.
  • Performed local feature selection using Recursive Feature Elimination with Cross-Validation.
  • Evaluated performance on three heterogeneous lung cancer datasets from Kaggle.
  • Achieved an accuracy of 91.17%.
  • AUC performance reached 0.9667, indicating high predictive capability.
  • Outperformed individual client models in terms of robustness and generalization.

Abstract

Aim/Purpose: Federated learning (FL) enables collaborative model training without centralizing raw data, thereby preserving privacy compared to traditional centralized machine learning approaches. To develop a privacy-preserving and explainable framework – where explainability is achieved through local SHAP analyses at each client and an aggregated global SHAP view – for health risk prediction in decentralized, heterogeneous environments, addressing the challenge of maintaining data confidentiality while achieving high predictive performance. Background: Traditional centralized health analytics solutions compromise data privacy; this paper introduces a federated learning approach that enables collaborative model building without exposing raw data. Methodology: The study integrates TreeNetX, a stable ensemble meta-learner, with SHAP-based explainability, leveraging client-specific training using gradient boosting models (XGBoost, LightGBM, CatBoost) and federated aggregation to address non-IID client distributions. Each client performs local feature selection using Recursive Feature Elimination with Cross-Validation (RFECV) before training. Experiments were conducted on three heterogeneous, tabular lung-cancer datasets from Kaggle to validate the model’s performance, using metrics including accuracy, precision, recall, F1-score, AUC, and calibration-related measures. Contribution: This paper contributes a novel federated ensemble learning framework that enhances predictive accuracy, robustness, and model interpretability, while ensuring data privacy in sensitive healthcare applications. Findings: The federated TreeNetX model achieved 91.17% accuracy and an AUC of 0.9667. It outperformed individual client models in robustness and generalization. SHAP-based analysis provided clinically meaningful insights, enhancing model trustworthiness. Recommendations for Practitioners: Practitioners should adopt federated ensemble strategies like TreeNetX for privacy-conscious predictive analytics to maintain compliance with data protection regulations while achieving high model performance. Recommendation for Researchers: Researchers should further explore hybrid federated architectures that combine explainability and advanced ensemble techniques to optimize interpretability and performance across heterogeneous environments. Impact on Society: The approach promotes ethical, privacy-preserving AI adoption in healthcare and other sensitive fields, contributing to safer and more trustworthy smart systems. Future Research: Future studies should investigate scaling the framework to even larger, more diverse client networks, integrating additional explainability tools, and extending applications beyond healthcare to other regulated industries.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shariff et al. (2025) studied this question.

synapsesocial.com/papers/69897a35f0ec2af6756e887chttps://doi.org/10.28945/5613
Ask AI
Helpful
Bookmark
Share
View Full Paper