Lung cancer can be discovered at an early stage to enhance patient survival. However, existing screening tools are both resource-intensive and inaccessible in low-resource countries. This paper introduces a machine learning model that uses an ensemble approach to predict lung cancer from a survey-based dataset of individuals based on symptoms. The suggested method leverages imbalanced data by using class-weighted learning and a stratified train-validation-test split to prevent data leakage and optimizing the decision threshold on the validation set to maximize clinical sensitivity. Several ensemble models were tested, and CatBoost achieved the best validation performance. The optimized model reached an accuracy and ROC-AUC of 95.16 and 93.75, respectively, on the held-out test set, with perfect recall and no false negatives. Extensive analyses, including calibration, subgroup analyses, performance analyses, feature importance analyses, and risk-stratification evidence, demonstrate the soundness and readability of the proposed framework. The above findings suggest that symptom-based ensemble learning models may be useful as supplementary measures for the initial risk evaluation and clinical triage of lung cancer.
Yousuf Nasser Al Husaini (Mon,) studied this question.