PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 16, 2026International Journal of Advances in Intelligent Informatics0 citationsOpen Access

Predict customer churn in the banking sector: a machine learning approach with imbalanced data handling techniques

JLJong-Hwa LeeVNVan-Ho NguyenHLHoanh-Su Le

Key Points

  • The study aims to predict customer churn in the banking sector by using machine learning techniques and addressing class imbalance in the dataset.
  • Analyzed a dataset of 10,127 customer records with 1,627 churned customers.
  • Applied data balancing methods like SMOTE and class-weight adjustments to handle imbalance.
  • Evaluated multiple machine learning models including Random Forest and Support Vector Machine for churn prediction.
  • Random Forest model achieved an 86% F1-score after employing SMOTE-Tomek Links.
  • Demonstrated strong predictive capability for customer churn in the banking sector.

Abstract

Customer value analysis is a critical component in formulating effective marketing and customer relationship management (CRM) strategies, especially in sectors where client movement and strong competition are prevalent A key element of this process lies in enhancing customer retention rates, as retaining existing clients is typically more cost-effective than acquiring new ones and directly contributes to improving overall profitability. In today’s banking environment, where customers can choose from a broad range of financial services, customer churn has become a critical challenge. Predicting and understanding attrition enables financial institutions to implement proactive and targeted interventions to protect market share and strengthen customer loyalty. This study analyzes a real-world dataset comprising 10,127 customer records from a commercial bank, where only 1,627 entries correspond to churned customers, thereby presenting a notable class imbalance problem. To address this, several data balancing techniques were applied, including class-weight adjustment, SMOTE, SMOTE-Tomek Links, and SMOTE-ENN. Multiple machine learning models - Support Vector Machine, Random Forest, Decision Tree, Logistic Regression, AdaBoost - were evaluated to identify the most effective approach for churn prediction. The Random Forest model achieved an 86% F1-score after applying SMOTE-Tomek Links, demonstrating strong predictive capability. The key contribution of this study lies in integrating advanced resampling techniques with ensemble learning and customer behavioral insights to improve churn prediction performance and support data-driven retention strategies in the banking sector.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lee et al. (2026) studied this question.

synapsesocial.com/papers/69b79d538166e15b153aab98https://doi.org/10.26555/ijain.v12i1.2262
Ask AI
Helpful
Bookmark
Share
View Full Paper