Abstract Introduction: Esophageal cancer is a significant health concern worldwide, accounting for 3.1% of all cancer burdens and 5.5% of all cancer-related deaths. Due to its impact, interest in adopting advanced methodologies has increased. Machine learning techniques offer a promising approach for gaining a deeper understanding of this disease. Methodology: The study is based on a case‒control study design, with a total of 400 case‒control subjects equally distributed. The study examined various machine learning-based prediction models, and for each model, several performance metrics, including accuracy, precision, F1 score, recall, and ROC-AUC, were evaluated. To optimize each model and determine the importance of the factors, a fivefold cross-validation technique was employed, and the ranking of feature importance was performed based on the weights in each model. Results: This study identified the Extra Tree Classifier model as the optimal approach for predicting esophageal cancer, with a model accuracy of 87.50%, a sensitivity of 92.5%, and a specificity of 80%. Compared with the top 10 risk factors on the basis of weight of feature importance, the model yielded an ROC-AUC value of 0.913, representing a substantial improvement of 10.1% over the baseline value of the traditional risk prediction model (ROC-AUC 0.812; 95% CI 0.59–0.94). Conclusion: The extra tree classifier model exhibited higher predictability and accuracy in identifying significant predictors of esophageal cancer. The incorporation of this machine learning-based model presents exciting opportunities for policymakers to focus on specific risk factors.
Kalita et al. (Mon,) studied this question.