PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 2026Journal of Chemical Theory and Computation0 citations

Hitchhikers Guide To Training More General Machine Learning Potentials in Heterogeneous Catalysis

View Full Paper
CWChenyu WuCYChao YangZFZhening Fang

Key Points

  • The research aims to enhance the training of machine learning potentials (MLPs) for more general applications in heterogeneous catalysis.
  • Developed REICO method to generate diverse, nonequilibrium training data.
  • Investigated distinct issues encountered during MLP training on diverse data sets.
  • Provided recommendations for data cleaning, model selection, and evaluation metrics for both performance and validation.
  • Identified critical deviations from standard training practices for system-specific MLPs.
  • Demonstrated the necessity of diverse training data for robustness in MLPs.
  • Enhanced understanding of how to better evaluate and select models for heterogeneous catalysis.

Abstract

Machine learning potentials (MLPs) have emerged as powerful simulation tools for heterogeneous catalysis. While current MD-active-learning workflows excel at fitting a specific system/reaction pathway through iterative structural sampling, the transition toward more general and transferable MLPs, designed to handle diverse structures and reactive events within a defined chemical space, presents fundamentally new challenges. Such generality often requires highly diverse, nonequilibrium training data, for which standard practices may confront challenges regarding training discipline and evaluation logic. Here, using our recently developed REICO method as an example to generate such data sets, we systematically investigate the distinct pathologies that arise when training on such diverse data, revealing critical deviations from standard system-specific MLP training. We further provide detailed recommendations on data cleaning, model selection, and error metrics for both numerical performance and physical validation, offering practical guidance for training MLPs with diverse and hybrid data sets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wu et al. (2026) studied this question.

synapsesocial.com/papers/698d6e5a5be6419ac0d5402fhttps://doi.org/10.1021/acs.jctc.5c02077
Ask AI
Helpful
Bookmark
Share
View Full Paper