Key points are not available for this paper at this time.
Advancement in technology has exponentially increased the number of digital transactions occurring every day. In order to make accurate decisions in a situation, we use various ML algorithms to identify hidden patterns in the data. However, the dynamic nature of a customer’s behavior and adaptive nature of fraudsters makes it difficult for accurate predictions by fraud detection systems. Additionally, the redactions (due to security reasons) in the publicly available credit card fraud datasets make it difficult for an in-depth analysis. This paper studies the various data imbalance handling techniques such as Oversampling, UnderSampling, Synthetic Minority OverSampling and Adasyn and its effects on the accuracy of the machine learning algorithms for credit card fraud detection. For an in-depth analysis, four machine learning algorithms like Logistic Regression, Random Forest, XGBoost and Decision Trees have been chosen for comparison. Each of these algorithms were tuned for optimal hyperparameters before building the models. As an additional feature, the benefits of feature engineering have been studied and a real-time fraud detection system has been developed and deployed on Google Cloud Platform. XGBoost with oversampling provided the best results with an ROC Score of 0.97 and is considered the best model. Therefore, the XGBoost model was deployed on Google Cloud and feature engineering was studied on this model. The model built with feature engineering had an ROC score of 0.9880 which is higher in comparison to the model without feature engineering 0.9659. Therefore, we can conclude that feature engineering significantly improves the accuracy of the model. Finally, a dashboard was built and deployed on Google Cloud Dashboards for incorporating analytics into the fraud detection system.
Attivilli et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: