PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 19, 2026Journal of Agricultural and Food Chemistry2 citations

Unlocking Enzyme Discovery: A High-Resolution Gene Cluster Database Powered by Phylogenetic Insights and Machine Learning

View Full Paper
SZSidun ZhangJHJunlong HeXDXuguo Duan

Key Points

  • This research aims to enhance enzyme discovery using a high-resolution phylogenetic database and machine learning techniques.
  • Developed an integrated high-resolution phylogenetic database across kingdoms.
  • Employed multilocus phylogeny to identify candidate enzymes.
  • Utilized a protein language model for predicting enzyme activities.
  • Implemented multilevel contact scoring to eliminate false positives.
  • Identified previously undocumented enzyme homologues including FadB, BktB, Ter, and YdiI.
  • Achieved R² value of 0.68 for the activity model, improving RMSE by 11% over previous methods.
  • Enhanced early enrichment (EF1%) by 16-fold through contact scoring.
  • Increased FadB yields from 0.65 g/L to 1.7 g/L in shake flasks, reaching 10.2 g/L in fermentation.

Abstract

High-throughput sequencing has generated vast genomic repositories that remain under-annotated, hampering enzyme discovery. We present an integrated pipeline that (i) builds a high-resolution, cross-kingdom phylogenetic database, (ii) mines candidates via multilocus phylogeny, (iii) predicts activities using an evolutionary-scale protein language model, and (iv) removes false positives through multilevel residue-atom contact rescoring. When applied to the r-BOX pathway, this approach uncovered numerous previously undocumented FadB, BktB, Ter, and YdiI homologues. Our activity model achieved R2 = 0.68 and reduced the RMSE on high-value targets by 11% compared to the prior SOTA (UniKP). Contact scoring improved early enrichment (EF1%) by 16-fold. Experimental validation targeting FadB increased titers from 0.65 g/L (shake flasks) to 1.7 g/L, reaching 10.2 g/L in a fermentation process. Together, these results establish a robust, generalizable framework for discovering scarce, high-value enzymes and prioritizing functional variants at scale.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/69bb91c7496e729e6297f338https://doi.org/10.1021/acs.jafc.5c06841
Ask AI
Helpful
Bookmark
Share
View Full Paper