Key points are not available for this paper at this time.
Abstract When conventional supervised machine learning (ML), including decision trees, support vector machines (SVMs), and k‐nearest neighbors (KNN), is applied to geological problems involving complex data sets, it is necessary to select a subset of raw or pre‐processed data types that will be used as input to the ML model. We revised four ML case studies involving 2D and 3D structural and lithological inference from geophysical survey data (magnetic, gravity, and radiometric measurements). For each study, we identified the most relevant inputs via a package of input selection approaches, comprising descriptive statistics, principal component analysis (PCA), correlation coefficient analysis, significance testing, and algorithmic input selection techniques, including Pearson, Spearman, Kendall, Minimum Redundancy Maximum Relevance (MRMR), Relief, permutation importance, Local Interpretable Model‐agnostic Explanations (LIME), and Shapley values. As anticipated, strategic input selection reduced collinearity among the raw input data sets and their standard derivatives, and consistently enhanced model performance across all case studies, improving accuracy, precision, and recall while reducing overfitting. However, different input selection methods proved optimal in different case studies, with no single approach consistently outperformed others across all geological contexts. This demonstrates the importance of using multiple complementary input selection methods when developing ML applications for automated geological mapping. When applying ML models to new localities with different geological contexts, data‐driven input selection and feature engineering should be revisited to ensure model performance, rather than assuming direct transferability of input configurations from the original study area.
Xu et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: