PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 27, 2026Infrastructures0 citationsOpen Access

Explainable Hybrid Intelligence for Predicting Tunnel Water Inrush Quantity Under Small-Sample, High-Heterogeneity Conditions: GAN Augmentation and Swarm-Optimized CatBoost

View Full Paper
RHRui HuangYCYuehua ChenLWLanjing Wang

Key Points

  • The study aims to develop an explainable model for predicting water inrush quantity in tunnels under challenging geological conditions.
  • Compiled a dataset with 55 real samples and 11 test samples using hydrogeological indicators.
  • Applied a generative adversarial network for data augmentation and used swarm optimization for CatBoost tuning.
  • Implemented multifaceted interpretability tools, including SHAP and ICE diagnostics, for model transparency.
  • Training-set augmentation led to improved accuracy for baseline models and the proposed hybrid model.
  • Identified lithologic and reflector-related factors as key influences on predictions with noted non-linear responses.
  • Reported accuracy based on limited samples should be seen as preliminary, needing further confirmation.

Abstract

This study aims to explore a leakage-aware and explainable machine learning framework for predicting tunnel water inrush quantity (WIQ) under small-sample and high-heterogeneity geological conditions. A project-level dataset was compiled at a fixed spatial granularity of 30 m per excavation segment by integrating forward prospecting outputs, construction-face observations, and geological reports, and six hydrogeological–structural indicators were used to predict the water inflow rate in cubic meters per hour. To overcome data scarcity and improve generalization, a tabular generative adversarial network (GAN) was introduced to augment the training distribution while preserving marginal statistics and inter-variable dependence, and a swarm-intelligence optimizer was employed to tune a Categorical Boosting (CatBoost) regressor for stable performance. In addition, six mainstream tree-based learners were benchmarked under a unified protocol, and model transparency was ensured through a multi-level interpretability suite combining SHapley Additive exPlanations (SHAP) attribution, partial dependence with individual conditional expectation (ICE) diagnostics, and interaction surfaces. Results show that, under the present fixed split, training-set augmentation was associated with improved performance for the evaluated baseline learners, and the proposed hybrid model achieved encouraging hold-out accuracy. However, because the dataset contains only 55 real samples and the test set contains only 11 real samples, the reported performance should be interpreted as an initial project-specific indication rather than robust evidence of generalizable reliability. Interpretability analyses further identify lithologic and reflector-related factors as dominant drivers, and reveal nonlinear response patterns and interaction-sensitive high-risk regions. Overall, the proposed framework shows potential to improve predictive performance and engineering interpretability for the studied project, and may provide a useful reference for drainage and reinforcement planning. Further confirmation through repeated data splitting, additional samples, and external validation is still needed before broader application.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Huang et al. (2026) studied this question.

synapsesocial.com/papers/6a168b160c924ddd1bd59f83https://doi.org/10.3390/infrastructures11060183
Ask AI
Helpful
Bookmark
Share
View Full Paper