PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 28, 2026Applied System Innovation3 citationsOpen Access

Explainable Hybrid CNN–XGBoost Framework for Multi-Class IoT Intrusion Detection with Leakage-Aware Feature Selection

DADeemah AlFuraihLMLotfi MhamdiAKAbdullah S. Karar

Key Points

  • To develop an explainable hybrid CNN–XGBoost framework for multi-class classification of IoT attacks.
  • Preprocessing network-traffic records using a chunk-wise workflow.
  • Feature selection based on Random Forest importance ranking.
  • Comparison of leakage-prone and leakage-aware feature ranking strategies.
  • Use of CNN for learning a compact feature representation.
  • Application of XGBoost for final multi-class classification.
  • Achieved 0.9324 accuracy under the leakage-aware protocol.
  • Obtained a macro-F1 score of 0.5910.
  • Leakage-aware selection provides a better estimate of generalization.
  • Identified that a small number of features heavily influence decisions.

Abstract

The rapid deployment of Internet of Things (IoT) devices has increased exposure to a diverse array of evolving cyberattacks, motivating the need for accurate and interpretable intrusion detection systems (IDS). In this work, we develop an explainable hybrid Convolutional Neural Network–Extreme Gradient Boosting (CNN–XGBoost) framework for multi-class IoT attack classification using the CIC IoT-DIAD 2024 dataset. Network-traffic records are preprocessed and standardized using a scalable, chunk-wise workflow, after which a compact top-k subset of features is selected via Random Forest importance ranking. To reduce selection bias, a leakage-prone feature-ranking strategy is compared with a leakage-aware strategy in which features are ranked using only the training data within each split. Subsequently, a one-dimensional Convolutional Neural Network (CNN) learns a 128-dimensional representation from the selected predictors, and XGBoost performs the final multi-class classification. Under the leakage-aware protocol, the proposed model achieves 0.9324 accuracy with 0.5910 macro-F1. Results indicate that leakage-aware selection provides a more defensible estimate of generalization while maintaining competitive detection performance. Finally, SHapley Additive exPlanations (SHAP) is used to interpret the model’s decisions in the learned latent space. The analysis shows that only a small number of embedding dimensions contribute most of the decision evidence, which can aid analyst triage, although the explanations remain indirect with respect to the original traffic features.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

AlFuraih et al. (2026) studied this question.

synapsesocial.com/papers/69a287690a974eb0d3c0317bhttps://doi.org/10.3390/asi9030049
Ask AI
Helpful
Bookmark
Share
View Full Paper