AgPd alloy catalysts frequently undergo structural evolution during catalytic reactions. Understanding the mechanisms driving this restructuring requires experimental in situ techniques, which are often expensive, as well as computationally intensive approaches such as machine learning potentials and molecular dynamics simulations. To address these challenges, this study develops a deep reinforcement learning (DRL) framework that integrates proximal policy optimization (PPO) with hierarchical hybrid action spaces and atomic-level state observations, enabling the prediction of surface dynamics beyond the reach of conventional methods. Compared to trust region policy optimization, the PPO-based DRL framework exhibits enhanced stability, accelerated convergence kinetics, and superior asymptotic performance. The DRL framework identifies a transition state with an energy barrier of 1.61 eV along the reconstruction pathway, while nudged elastic band calculations yield a comparable barrier of 1.59 eV. Notably, the DRL framework successfully identifies six distinct surface configurations in AgPd alloys, surpassing the four configurations found using traditional minima hopping methods. Analysis reveals that the final reconstructed surface undergoes facet slip accompanied by atomic substitution, and energy barrier during reconstruction pathway is 1.51 eV at 600 K. The DRL framework uncovers a series of Pd–Ag swap actions wherein Ag preferentially migrates to the surface, driven by surface energy minimization (with Ag surface energy of 1.246 J/m2 vs Pd surface energy of 2.003 J/m2) and strain release. This work represents the first application of a PPO-based DRL approach to predict AgPd surface reconstruction dynamics.
Wei et al. (Thu,) studied this question.