PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 26, 2026Journal of Chemical Information and Modeling7 citations

PeptideNet: An Integrative Deep Learning Framework for Predicting Diverse Bioactive Peptides Using Protein Language Model Embeddings

View Full Paper
HZHamza ZahidMMaryamKCKil To Chong

Key Points

  • The aim is to develop an effective model for predicting the bioactivity of diverse bioactive peptides.
  • Investigated five categories of bioactive peptides using four feature representations.
  • Developed 20 hybrid deep learning models combining CNNs and BiGRUs.
  • Utilized large protein language model embeddings and physicochemical descriptors.
  • Achieved accuracies of 0.83 to 0.94 for various peptide types on independent data.
  • ESM-2 embeddings outperformed other feature sets across all peptide categories.
  • t-SNE visualization and sequence logo analysis confirmed effective representation and conserved patterns.

Abstract

Bioactive peptides are multifunctional biomolecules composed of short amino acid sequences that exhibit diverse biological activities, including antioxidative, antihemolytic, anticell-penetrating, antiviral, and antimicrobial effects. Accurate computational prediction of peptide bioactivity is essential for accelerating the discovery and design of peptide-based therapeutics. In this study, we investigated five categories of bioactive peptides using four distinct feature representations, including large protein language model embeddings (ESM1, ESM2, and ProtBert) and physicochemical descriptors. A total of 20 hybrid deep learning models integrating Convolutional Neural Networks (CNNs) and Bidirectional Gated Recurrent Units (BiGRUs) were developed to capture both local sequence motifs and long-range dependencies. The proposed PeptideNet model achieved robust predictive performance, with accuracies of 0.83, 0.87, 0.89, 0.92, and 0.94 for antioxidative, antihemolytic, anticell-penetrating, antiviral, and antimicrobial peptides, respectively, on the independent data set. Among the evaluated feature sets, ESM-2 embeddings consistently outperformed others across all peptide types, providing rich contextual and evolutionary information. Furthermore, t-SNE visualization of learned representations demonstrated effective generalization across peptide classes, while positional sequence logo analysis revealed conserved residue patterns contributing to peptide bioactivity. The integration of large protein language model embeddings with the PeptideNet architecture enables the model to capture both global contextual information and residue-level features, establishing a generalized and interpretable framework for multipeptide bioactivity prediction.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zahid et al. (2026) studied this question.

synapsesocial.com/papers/699fe39d95ddcd3a253e7988https://doi.org/10.1021/acs.jcim.5c02885
Ask AI
Helpful
Bookmark
Share
View Full Paper