Summary Carbonate lithology identification has long posed challenges in petroleum exploration due to overlapping well logging responses caused by compositional similarities and the thin-bedded nature of transitional lithologies. Although machine learning (ML) has provided new approaches to address this issue, the “black-box” nature of most ML models limits their geological interpretability and practical reliability. In this study, a high-precision lithology classification model was applied and coupled with the Shapley additive explanations (SHAP) framework to quantitatively evaluate the contribution and directionality of each logging feature to the model’s predictions. By combining K-means oversampling with the XGBoost algorithm, the training data were effectively balanced, and the robustness of the model was improved, ultimately achieving an overall classification accuracy of 92%. SHAP analysis further clarified the controlling influence of key logging parameters on lithology discrimination, revealing the model’s internal decision logic in a physically meaningful manner. This research establishes an interpretable and data-driven workflow for carbonate lithology identification and provides a theoretical and technical reference for the reliable application of ML in complex carbonate reservoir characterization.
Li et al. (Fri,) studied this question.