Lung cancer is still one of the major contributors to cancer-related mortality. Treatment effectiveness depends directly on early diagnosis and accurate prediction of disease progression. This study focuses on the application of classical machine learning algorithms —including Logistic Regression, Support Vector Machine, Random Forest, and XGBoost —for questionnaire-based lung cancer risk assessment. A comparative study of six machine learning algorithms used to predict lung cancer based on survey data was conducted. According to the ROC-AUC metric, Random Forest model achieved the best result. In terms of F1-score and precision for the positive class, the support vector machine (SVM) demonstrated the highest efficiency. However, gradient boosting provided the most balanced results across key clinical metrics. SHAP (SHapley Additive exPlanations) analysis identified three preliminary predictors potentially associated with lung cancer risk in this sample: seeing a pulmonologist, unexplained weight loss, and loss of appetite. These preliminary results are consistent with clinical observations and suggest the potential interpretability of ensemble machine learning approaches in medical diagnostics, though confirmation on larger datasets is required. Overall, the preliminary results suggest the potential of using machine learning methods for lung cancer screening based on questionnaire data, particularly with an expanded training sample. Potential applications of early lung cancer diagnosis are discussed. Further research in this area will be conducted in conjunction with image recognition of computed tomography scans.
Карымсакова et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: