PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 8, 2026Big Data and Cognitive Computing2 citationsOpen Access

Predicting Bond Defaults in China: A Double-Ensemble Model Leveraging SMOTE for Class Imbalance

View Full Paper
CTChongwen TianLRLi Rong

Key Points

  • The study aims to improve financial bond default prediction by addressing severe class imbalance using a new model.
  • Introduces the Double-Ensemble Learning Classification with SMOTE (DELC-SMOTE) framework.
  • Utilizes introspective stacking with six improved base learners via meta-learning.
  • Implements performance-weighted voting for fusing optimized classifiers.
  • Achieves a G-mean of 0.9152 and a Specificity of 0.8715.
  • Model outperforms standard classifiers across various imbalance ratios.
  • Demonstrates strong performance against data noise and outliers.

Abstract

This study proposes the Double-Ensemble Learning Classification with SMOTE (DELC-SMOTE), a novel hierarchical framework designed to address the critical challenge of severe class imbalance in financial bond default prediction. The model integrates the Synthetic Minority Over-sampling Technique (SMOTE) into a two-phase ensemble architecture. The first phase employs introspective stacking, where six heterogeneous base learners are individually enhanced through algorithm-specific balancing and meta-learning. The second phase fuses these optimized experts via performance-weighted voting. Empirical analysis utilizes a comprehensive dataset of 10,440 Chinese corporate bonds (522 defaults, ~5% default rate) sourced from Wind and CSMAR databases. Given the high cost of both false negatives and false positives in risk assessment, the Geometric Mean (G-mean) and Specificity are employed as primary evaluation metrics. Results demonstrate that the proposed DELC-SMOTE model significantly outperforms individual base classifiers and benchmark ensemble variants, achieving a G-mean of 0.9152 and a Specificity of 0.8715 under the primary experimental setting. The model exhibits robust performance across varying imbalance ratios (2%, 10%, 20%) and strong resilience against data noise, perturbations, and outliers. These findings indicate that the synergistic integration of data-level resampling within a diversified, two-tiered ensemble structure effectively mitigates class imbalance bias and enhances predictive reliability. The framework offers a robust and generalizable tool for actionable default risk assessment in imbalanced financial datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tian et al. (2026) studied this question.

synapsesocial.com/papers/69acc5bd32b0ef16a405065chttps://doi.org/10.3390/bdcc10030081
Ask AI
Helpful
Bookmark
Share
View Full Paper