Abstract Heterogeneous catalysis optimization relies on the coordinated selection of catalyst composition, synthesis procedure, and reaction conditions, yet the empirical links between these variables and catalytic performance remain scattered across individual studies. Here, we show that an interpretable machine learning (ML) framework can decode this complexity by integrating disparate literature data. We curate, to our knowledge, the first unified dataset from 200 papers spanning three decades of CO 2 -to-methanol synthesis over Cu/ZnO-based catalysts, and derive a subset of apparent activation energy ( E app ). Using the dataset, we optimize six ML models via Bayesian hyperparameter tuning and apply regularization strategies to mitigate overfitting. We then use SHAP to quantify the non-linear impact of catalyst properties, synthesis procedures, and reaction conditions on methanol selectivity, CO 2 conversion, and methanol space-time yield. Crucially, this data-driven approach uncovers previously obscured operating windows for achieving high methanol productivity at milder pressures than conventionally assumed. Extending the ML framework to the E app kinetic subset further identifies regime-sensitive signatures. A controlled experiment over commercial catalysts corroborates the positive space velocity– E app relationship inferred from ML models. These results show how interpretable ML built on curated legacy experiments can support operating-window selection, hypothesis-driven kinetic experiments, and practical optimization of mature catalytic technologies.
Yan et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: