PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 25, 2026PLoS Computational Biology0 citationsOpen Access

A novel transformer-based platform for the prediction and design of biosynthetic gene clusters for (un)natural products

View Full Paper
TKTomoki KawanoTSTaro ShiraishiTKTomohisa Kuzuyama

Key Points

  • The aim is to develop a transformer-based platform for accurate prediction and design of biosynthetic gene clusters responsible for natural products.
  • Developed a transformer-based framework using RoBERTa architecture
  • Trained on datasets including bacterial and fungal genomes
  • Evaluated using 2,492 experimentally-validated biosynthetic gene clusters from MIBiG
  • Compared predictions of BGC-trained and genome-trained models on a specific BGC case study
  • Over 50% of true domains were ranked first in predictions
  • More than 75% of true domains appeared within the top 10 predicted candidates
  • Achieved over 70% classification accuracy for major compound classes like polyketides and terpenes
  • Identified additional predicted domains not present in the original BGC

Abstract

Biosynthetic gene clusters (BGCs), comprising sets of functionally related genes responsible for synthesizing complex natural products, are a rich source of bioactive compounds with pharmaceutical potential. Here, we present a transformer-based framework that models functional domains as linguistic units to capture and predict their positional relationships within genomes. Using a RoBERTa architecture, we trained models on four progressively broader datasets: bacterial BGCs, Actinomycetes genomes, bacterial genomes, and bacterial plus fungal genomes. Evaluation using 2,492 experimentally-validated BGCs from the MIBiG database showed that more than 50% of true domains were ranked first and over 75% within the top 10 candidates. Our models also achieved classification accuracies exceeding 70% for major compound classes including polyketides (PKs) and terpenes. To explore model-guided BGC design, we compared predictions from the BGC-trained and genome-trained models using the BGC for the bacterial diterpenoid cyclooctatin as a case study. The genome-trained model uniquely predicted several domains absent from both the original BGC and the prediction by the BGC-trained model. Heterologous expression of one of those predicted domains in Streptomyces albus , together with the biosynthetic genes for cyclooctatin, yielded an unknown cyclooctatin derivative. This framework not only provides a novel BGC prediction method using machine learning but also facilitates rational design of artificial BGCs. Future integration of transcriptomic, protein structural, and phylogenetic data will enhance the models’ predictive and generative capabilities, supporting accelerated discovery and engineering of natural products.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kawano et al. (2026) studied this question.

synapsesocial.com/papers/699e91fdf5123be5ed04fd23https://doi.org/10.1371/journal.pcbi.1013181
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A Cluster Separation Measure1979 · 9,068 citations
  2. 2Silhouettes: A graphical aid to the interpretation and validation of cluster analysis1987 · 21,514 citations
  3. 3A dendrite method for cluster analysis1974 · 6,822 citations
  4. 4Introduction to Information Retrieval2008 · 11,120 citations
  5. 5Regulation of secondary metabolism in streptomycetes2005 · 831 citations