PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 21, 2026Systems2 citationsOpen Access

An Interpretable Credit Default Risk Prediction Framework Integrating Causal Feature Selection and Double Machine Learning

View Full Paper
TCTinggui ChenRZRui ZhangJHJian Hou

Key Points

  • The aim is to develop a predictive model for credit default risk that integrates causal feature selection and enhances interpretability.
  • Utilized information value and extreme gradient boosting for initial feature reduction.
  • Employs the Peter Clark algorithm with perturbation and bootstrap sampling for stable causal feature identification.
  • Constructs higher-order interaction features to improve nonlinear modeling.
  • Integrates causal forest double machine learning to estimate causal effects and includes a counterfactual explanation mechanism.
  • Causal features improved model performance, reflected in higher F1 scores, AUC, and G-mean.
  • Logistic regression models showed significant gains due to the addition of causal features.
  • The new framework enhances interpretability, stability, and compliance in high-risk financial scenarios.

Abstract

In the context of the rapid advancement of financial technology, the issue of credit card default has become increasingly salient, emerging as one of the crucial risks that financial institutions are eagerly addressing. Traditional credit card default risk prediction models predominantly rely on statistical correlations for feature selection. This approach not only makes it challenging to uncover the genuine causal relationships between variables but also leads to limitations in prediction accuracy and interpretability. To overcome these limitations, this paper presents a novel credit card default risk prediction model that integrates causal feature screening, interaction feature construction, and interpretability enhancement. Initially, by leveraging the information value (IV) and eXtreme gradient boosting (XGBoost), we perform initial feature dimensionality reduction. Subsequently, we introduce the Peter Clark algorithm (PC) augmented with perturbation enhancement and bootstrap sampling to identify a stable set of causal features. Building on this foundation, we proceed to construct higher-order interaction features to bolster the model’s nonlinear modeling capacity. These causal features and their interaction counterparts are then fed into a variety of mainstream machine learning models for training and evaluation purposes. Furthermore, on the basis of the causal feature set identified via the PC algorithm, we construct a causal path diagram. We also incorporate the causal forest double machine learning (causal forest DML) method to estimate the causal effects of features. Additionally, we design a counterfactual explanation mechanism to aid in analyzing the direction and magnitude of the impact of variable interventions on default probability. Empirical tests conducted using four typical credit datasets reveal the following findings: (1) the introduction of causal features generally enhances the model’s performance in terms of the F1 score, area under the curve (AUC), and geometric mean (G-mean). This improvement is especially pronounced in models that are highly reliant on feature quality, such as logistic regression (LR). (2) Causal features offer significant advantages in terms of model interpretability, stability, and compliance, thereby presenting a new research paradigm for credit risk prevention and control in high-risk financial scenarios.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chen et al. (2026) studied this question.

synapsesocial.com/papers/69be37506e48c4981c676e1fhttps://doi.org/10.3390/systems14030327
Ask AI
Helpful
Bookmark
Share
View Full Paper