The rapid expansion of Internet of Things (IoT) infrastructures in smart city environments has increased the demand for reliable intrusion detection systems (IDS). However, many existing studies rely on single-dataset evaluations and inconsistent experimental settings, which can lead to overly optimistic performance estimates. In this study, we propose a standardized benchmarking framework for evaluating artificial intelligence-based IDS across heterogeneous IoT datasets, including CIC-IoT 2023, BoT-IoT, and N-BaIoT. Multiple classical machine learning and deep learning models are evaluated under a unified preprocessing pipeline and a consistent evaluation protocol. A hybrid CNN–BiLSTM–Attention architecture is also implemented as a reference model within this framework. While several models achieve near-perfect performance under intra-dataset evaluation, cross-dataset experiments reveal substantial performance degradation and unstable metric behavior under distribution shifts. These results highlight the limitations of dataset-specific optimization and emphasize the necessity of cross-dataset validation for realistic IoT intrusion detection evaluation. All experiments are conducted under a binary intrusion detection setting (benign vs. attack) to enable consistent comparison across datasets. Consequently, the reported results reflect binary detection performance and do not capture attack-type discrimination.
Alghamdi et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: