Hematological disorders are clinically common and frequently coexist within the same patient, yet most existing studies have adopted single-label classification frameworks that do not reflect this multi-morbid clinical reality. In this study, the performance of machine learning and deep learning algorithms was comparatively evaluated for the multi-label classification of 15 hematological conditions using complete blood count (CBC) parameters. Fifteen conditions were selected to encompass the most prevalent erythrocyte, leukocyte, and platelet line abnormalities encountered in routine practice, thereby reflecting the heterogeneous diagnostic landscape clinicians face when interpreting hemogram results. A retrospective dataset comprising CBC records of 88,798 patients, collected between 2022 and 2024 from Kars Harakani State Hospital and labeled by specialist physicians, was used. Six models—Random Forest, Logistic Regression, Linear SVM, Deep MLP, 1D-CNN, and Wide & Deep Network—were evaluated across 21 hemogram parameters. The dataset was partitioned into 80% training and 20% testing subsets, and class weighting strategies were applied to address class imbalance. Random Forest achieved the highest overall performance (F1-Micro: 0.9945, ROC-AUC: 0.9999, Subset Accuracy: 98.80%), followed by deep learning architectures. These results demonstrate that Random Forest is highly effective for multi-label hematological classification and holds considerable promise for integration into clinical decision support systems. With its large-scale dataset and comprehensive multi-label classification framework, this study represents one of the most extensive investigations of automated hematological disease diagnosis to date.
Yaman et al. (Tue,) studied this question.