PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 29, 2026Computers, materials & continua/Computers, materials & continua (Print)0 citationsOpen Access

Hierarchical Mixed-Effects and Stacked Machine Learning Ensembles with Data Augmentation for Leakage-Safe E-Waste Forecasting

HMHatim MadkhaliASAbdullah SheneamerLNLinh Nguyen

Key Points

  • The aim is to forecast e-waste generation in European countries using advanced statistical and machine learning techniques.
  • Used sparse panel data from 32 European countries spanning from 2005 to 2018.
  • Applied various forecasting approaches including ARIMA, LSTM, and hierarchical mixed-effects models.
  • Implemented data augmentation through bootstrapping to enhance model stability and accuracy.
  • Conducted a feature ablation study to identify critical features for forecasting, particularly lagged values.
  • Achieved a weighted validation R2 of 0.992 using stacking methods for the 2017–2018 period.
  • Time-series approaches showed negligible predictive power with mean R2 of -9683.
  • Bootstrapping reduced forecast RMSE by 18.6% for high-variance tonnage forecasts.
  • Feature analysis indicated only a few lags were necessary, reducing the risk of data leakage.

Abstract

Consumer electronics, with 62 million tons of electronic waste (e-waste) generated in 2022 and e-waste expected to grow to 82 million tons annually by 2030, pose critical challenges when it comes to national infrastructure and circular economy policies. This paper compares forecasting approaches using sparse panel data for 32 European countries (2005–2018, Eurostat/Waste Electrical and Electronic Equipment (WEEE) Directive), focusing on leakage-safe prospective validation to guarantee true predictive performance. We make one-step-ahead predictions with conservative features (primarily lagged values) to account for temporal autocorrelation but with reduced multicollinearity (Variance Inflation Factor (VIF) ≈ 1.0). Cross-paradigm comparisons such as time-series baselines Autoregressive Integrated Moving Average (ARIMA), Seasonal ARIMA (SARIMA), Long Short-Term Memory (LSTM), hierarchical mixed-effects models, pooled machine learning (9 methods), and block-bootstrap-augmented stacking ensembles demonstrate stacking’s effectiveness, with a weighted validation R2 of 0.992 for held-out 2017–2018 data. Time-series approaches demonstrate negligible predictive power (mean R2 = −9683) given non-stationarity and limited samples, while the hierarchical approach provides virtually no benefit (Intraclass Correlation Coefficient (ICC) 0.011) amidst computational instability. Bootstrapping improves high-variance tonnage forecasts (Root Mean Squared Error (RMSE) reductions of 18.6%) while being detrimental to stable units, thus reinforcing parsimony. Feature ablation validates that only a few lags are necessary, preventing leakage from rolling means or calendar trends. Our method enables conservative year-ahead forecasts with quantified uncertainty, conservative estimates of e-waste management, allowing for buffer planning and policies even when data is scarce. By using strictly out-of-sample tests rather than biased ones, this work characterizes achievable year-ahead performance under sparse annual panels.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Madkhali et al. (2026) studied this question.

synapsesocial.com/papers/69c8c336de0f0f753b39dda7https://doi.org/10.32604/cmc.2026.074444
Ask AI
Helpful
Bookmark
Share
View Full Paper