Data-driven approaches are pivotal for optimizing single-atom catalysts. However, when the training data size is limited, the prevalence of multicollinearity among physicochemical features frequently destabilizes statistical models and obscures chemical interpretation. To address this challenge, we developed a wrapper-based feature selection algorithm termed adaptive reduction of multicollinearity with retention (ARMR). This method iteratively identifies feature subsets that satisfy strict multicollinearity constraints while minimizing cross-validated prediction error. The features selected by ARMR yielded robust predictive performance in both linear and nonlinear models, clearly outperforming baselines derived from univariate correlations or standard importance rankings. The efficacy of ARMR was validated through machine learning predictions using a density functional theory dataset of CO adsorption energies on 11 transition metals embedded as a single atom on three TiO2 polymorphs (rutile, anatase, and brookite). Specifically, a gradient boosting regression model trained on ARMR-selected features achieved high predictive accuracy with a test root-mean-square error of 0.087 eV. Furthermore, analysis of selected features identified s-orbital charge density as an important electronic feature correlating with CO binding strength. By removing multicollinearity, ARMR eliminates redundant properties intrinsically correlated with metal-specific characteristics like oxidation state, retaining only robust features. ARMR provides valuable insights into metal–support interactions and offers an effective strategy for accelerating rational design of single-atom catalysts through interpretable structure–property relationships.
Sawabe et al. (Tue,) studied this question.