We present a deep-learning framework that integrates substrate information for genome-scale mining of natural product biosynthetic enzymes. Biosynthetic enzyme mining is essential for promoting the heterologous expression and large-scale production of bioactive natural products. However, existing sequence-based approaches rely heavily on large amounts of homologous enzyme sequence data, which limits their accuracy and generalization capability. To address this limitation, ConESI predicts enzyme–substrate interactions directly from genomic candidate protein sequences, enabling the identification of functional enzymes associated with specific small-molecule substrates at the genome scale. We present ConESI, a deep-learning framework that integrates substrate information to mine natural product biosynthetic enzymes at the genome scale. By predicting enzyme–substrate interactions directly from genomic sequences, ConESI accurately identifies functional enzymes for known small-molecule substrates, overcoming the limited training data of sequence-based methods. We built a benchmark of 246 substrate–enzyme–genome triplets, covering 118 substrates, 236 enzymes, 206 genomes, and 2,137,735 non-redundant proteins. ConESI was evaluated in two tasks: predicting enzymes for a single substrate across 129 genomes and identifying enzymes for 118 substrates within native genomes. Compared with baseline methods, it achieved more stable performance in cross-species enzyme mining tasks (for the substrate C00035, the mean and median normalized positions were 9.58% and 8.53%, respectively, and the mean and median BEDROC 10 values were 41.80% and 42.63%, respectively). By introducing a metric-discriminative contrastive learning strategy, ConESI improves both cross-substrate and cross-species mining accuracy and stability in enzyme mining tasks, thereby enabling reliable functional annotation of biosynthetic enzymes and the discovery of cryptic metabolic pathways. Scientific contribution: We introduce ConESI, a general method for enzyme mining that explicitly integrates substrate information to predict enzyme–substrate interactions directly from annotated protein sequences in genomes, addressing cross-enzyme limitation of traditional sequence-only enzyme mining approaches. A metric-discriminative contrastive learning strategy was introduced to enhance the effectiveness of multimodal data fusion in enzyme–substrate interaction prediction. We establish a rigorously designed benchmark comprising 246 substrate–enzyme–genome triplets and more than two million protein sequences. By consistently outperforming baseline methods in both cross-species and cross-substrate tasks, ConESI demonstrates that introducing a metric-discriminative contrastive learning strategy can fundamentally improve enzyme mining accuracy and stability across species and substrate backgrounds.
No takes yet. Share an insight, caveat, or question.
Li et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: