PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 2026International Journal of Artificial Intelligence Tools2 citations

Beyond Accuracy: A Comprehensive Comparative Study of Gradient Boosting versus Tabular Deep Learning and Explainability Techniques for Mixed-Type Tabular Data Models Using SHAP and LIME

View Full Paper
ALAlina LazarPPPrativa PokhrelSDShrabony Rany Das

Key Points

  • Evaluate the performance of gradient boosting and deep learning models on varied tabular datasets.
  • Analyzed six diverse datasets including census and fraud data.
  • Compared gradient boosting models XGBoost and LightGBM with deep learning models like MLP and TabPFN.
  • Used performance metrics like accuracy, F1 score, ROC-AUC, and log loss for evaluation.
  • Applied SHAP and LIME for model interpretability and assessed explanation quality.
  • Gradient boosting models outperformed deep learning models in most metrics.
  • Gradient boosting maintained better calibration and interpretability across datasets.
  • SHAP insights highlighted the stability and clarity of feature attributions in gradient boosting.

Abstract

The goal of this study was to evaluate the performance of traditional gradient boosting (GB) and neural network models on diverse tabular datasets that differ in scale, class balance, and feature composition (numerical, categorical, or mixed). We focused on six representative datasets: adult census income, bank marketing, credit card fraud, breast cancer diagnosis, diabetes, and in-vehicle coupon recommendation, each with distinct challenges related to dimensionality, sample size, and heterogeneity. We benchmark the predictive performance of XGBoost and LightGBM (gradient boosting models) against Multilayer Perceptrons (MLP), Tabular Transformers, and tabular prior-data fitted network (TabPFN), using metrics such as accuracy, F1 score, ROC-AUC, and log loss. To ensure transparency and interpretability, we applied SHapley Additive exPlanations (SHAP) and Local Interpretable Model-Agnostic Explanation (LIME) to all models and evaluated the explanation quality using stability, fidelity, and consistency criteria. Our findings confirm that gradient boosting models consistently achieve the best balance of performance, calibration, and interpretability across heterogeneous and imbalanced datasets. SHAP-based insights show that gradient boosting (GB) models provide more stable and interpretable feature attributions, making them well suited for high-stakes domains such as finance and healthcare. These results emphasize the practical advantages of gradient boosting methods for structured data tasks and highlight the interpretability limitations of deep learning models when applied to tabular datasets. Future work will explore hybrid architectures and pretraining strategies to close this performance gap.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lazar et al. (2026) studied this question.

synapsesocial.com/papers/698d6d695be6419ac0d5257ehttps://doi.org/10.1142/s0218213026400038
Ask AI
Helpful
Bookmark
Share
View Full Paper