Empirical evaluation demonstrates proxy-driven disparities in credit scoring models, indicating that removing sensitive features fails to eliminate bias.
AI-driven credit scoring is supervised as a high-risk application in banking and insurance, yet unfairness is rarely operationalized as a measurable category of model, conduct, legal, and reputational risk. Using 20,000 anonymized applications from a Southern European digital lender (15.2% twelve-month default rate), we estimate three model families—a regularized logistic regression, a gradient-boosting machine, and a multi-layer perceptron—under a fully crossed design in which each family is evaluated without mitigation and under pre-processing (reweighing), in-processing (an exponentiated-gradient reduction, applicable to any base learner, together with adversarial debiasing where gradient-based training permits it), and post-processing (reject-option) interventions, so that the mitigation effect is no longer confounded with the choice of estimator. No sensitive-group field enters any estimated specification; group membership is used exclusively for auditing. Predictive performance (AUC-ROC, Brier score and Brier skill score relative to the base-rate forecast, F1 on the default class, Gini, and the Kolmogorov–Smirnov statistic) is reported jointly with group fairness (demographic-parity and equal-opportunity differences, disparate-impact ratio, Theil index) and with group-conditional calibration, at an explicitly stated and economically justified decision threshold. Every fairness quantity is accompanied by stratified-bootstrap confidence intervals and, for stochastic learners, by seed-level dispersion. The interpretable benchmark attains an AUC of 0.780 and a Brier score of 0.104 against 0.129 for the constant base-rate forecast, and the high-capacity models improve on it by under one AUC point. Disparity is present but is located geographically rather than in the composite group label: the disparate-impact ratio is 0.724 [0.693, 0.754] for the lowest socio-economic neighborhood cluster, excluding the four-fifths screening value, against 0.809 [0.776, 0.840] for the ethno-socioeconomic proxy, whose interval contains it, and no measurable gender disparity. Group membership is recoverable from the neutral feature set at an AUC of 0.654, and 42% of the group gap in predicted risk travels through the bureau credit score alone, so feature deletion cannot close the channel. Feature attributions and an auxiliary group-recoverability test locate the proxy pathways through which disparity arises, and a misclassification-sensitivity analysis bounds the effect of error in the group proxy, which attenuates measured disparity toward parity. We map the results onto Regulation (EU) 2024/1689 as amended by Regulation (EU) 2026/1744, the GDPR as interpreted in SCHUFA Holding, Directive (EU) 2023/2225, EBA loan-origination guidance, and Solvency II, EIOPA, and IAIS expectations, and propose fairness-risk controls organized around impact assessment, independent validation, and three lines of defense governance. Because the evidence comes from credit origination at a single lender, the insurance argument is developed at the level of regulatory and governance architecture rather than as an empirical transfer of estimates.
No takes yet. Share an insight, caveat, or question.
Paulo Alcarva (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: