The discovery of highly active polyethylene (PE) catalysts demands a systematic understanding of structure–condition–activity relationships in a vast chemical space. In this Letter, we present a data-driven framework combining explainable machine learning (ML) with large-scale virtual library generation. From a curated data set of 507 catalysts (bis (phenoxyimine) and bis (imino) pyridine ligands, seven metals), a gradient boosting regression (GBR) model achieves a test R 2 of 0. 91, outperforming convolutional and graph neural networks. SHAP analysis identifies topological (Chi2v), electronic (EStateVSA), and hydrophobic (SlogPVSA) descriptors as governing activity and reveals a classical volcano-type temperature dependence, fundamentally governed by the Sabatier principle. A virtual library of 665 685 structures, constructed via combinatorial fragment assembly, extends the known chemical space substantially. High-throughput screening, coupled with SCscore filtering, yields 1090 synthetically accessible candidates with predicted activities exceeding 2 × 10 7 g mol –1 h –1. Substructure analysis uncovers metal-dependent design rules, in which early transition metals favor electron-deficient aromatics while late metals profit from moderately sized alkyls. This work establishes a practical route from experimental data to actionable catalyst designs.
Li et al. (Sun,) studied this question.