With the broad application of machine learning (ML) and deep learning (DL) models in financial markets, it has become increasingly important to evaluate their performance compared with traditional methods, especially for credit risk scoring in financial services. This mechanism determines the probability of default (PD) for borrowers, in other words, how likely a client is to fail to repay their debt. With the rapid development and growing availability of ML/DL technologies, it is essential for banks, financial institutions, auditors and regulatory bodies to assess whether these approaches truly outperform traditional models in terms of accuracy, stability and interpretability or whether their complexity comes at a cost. A systematic review following PRISMA examined credit risk scoring models. From 520 initial articles, 117 were analyzed to compare ML/DL approaches with traditional methods. This review evaluates accuracy, stability and interpretability, offering guidance for model selection in real-world credit scoring. Logistic regression remains essential in regulated contexts requiring transparency, supporting informed decisions on balancing performance and explainability.
Aboulmaouda et al. (Tue,) studied this question.