The growing reliance on solar energy necessitates robust and interpretable forecasting models for stable grid management. Current research frequently employs Explainable AI (XAI) to glean insights from complex black-box models, yet the reliability and consistency of these explanations remain largely unvalidated. Inconsistent feature attributions can mislead grid operators by incorrectly identifying the dominant drivers of solar generation, thereby affecting operational planning, reserve allocation, and trust in AI-assisted decision-making. This study addresses this critical gap by conducting a systematic statistical evaluation of feature rankings generated by multiple XAI methods, including model-agnostic (SHAP, PDP, PFI, ALE) and model-specific (Split- and Gain-based) techniques, within a time-series regression context. Using a LightGBM model for one-day-ahead solar power forecasting across four sites in Calgary, Canada, we evaluate consensus and divergence using the Friedman test, Kendall’s W, and Spearman’s rank correlation. To ensure the generalizability of our findings, we further validate the results using a CatBoost model. Our results show a strong overall agreement across methods (Kendall’s W: 0.90–0.94), with no statistically significant difference in ranking (p > 0.05). However, pairwise analysis reveals that the “Split” method frequently diverges from other techniques, exhibiting lower correlation scores. These findings suggest that while XAI consensus is high, relying on a single method—particularly the split count—poses risks. We recommend employing multi-method XAI and using agreement as an explicit diagnostic to ensure transparent and reliable solar energy predictions.
Thinn et al. (Wed,) studied this question.