Macroeconomic indicators and their transformations are critical in credit risk modeling but often suffer from high multicollinearity, distorting model estimates and reducing predictive accuracy.Since these indicators are collected at varying frequencies, such as 3-month, 6-month intervals, transformations with lags and differences are required to align them, resulting in multiple derived features from the same indicator.Traditional methods like Principal Component Regression (PCR) reduce multicollinearity but compromise interpretability by converting features into latent components.To address this, we propose a hybrid feature selection framework that combines clustering algorithms (K-means, Agglomerative, DBSCAN) with Variance Inflation Factor (VIF) filtering.This approach preserves the original economic meaning of features-such as GDP growth, exchange rate changes, and inflation-while reducing dimensionality and multicollinearity.Experimental results show that the K-means + VIF method achieves best performance out of our experiments, with a test R-squared of 0.95653, MSE of 0.00395, RMSE of 0.06288, and a maximum VIF of 4.08053.These metrics demonstrate both high predictive accuracy and low multicollinearity.By retaining interpretable features and validating across pre-and post-pandemic periods, our framework offers a transparent solution for macroeconomic-based loan losses prediction.
No takes yet. Share an insight, caveat, or question.
A 2025 study studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: