PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 18, 2026IEEE Transactions on Computational Biology and Bioinformatics0 citations

A Comparative Study of Machine Learning Models for Identification of Antiviral Peptides Using Various Encoded Features

View Full Paper
MHMd. Zahid HasanSSShahriar ShakilTKTasmin Karim

Key Points

  • This research aims to evaluate machine learning models for identifying antiviral peptides from protein sequences using different feature encodings.
  • Introduced a light gradient boosting machine (LGBM) model with various feature encoding methods.
  • Evaluated eight machine learning algorithms using both single-feature and combined feature encodings.
  • Measured performance based on accuracy, precision, recall, F1-score, and AUC.
  • LGBM model achieved an accuracy of 98%, precision of 97%, recall of 98%, F1-score of 98%, and AUC of 1.00.
  • Combined feature encoding improved accuracy by approximately 3% compared to individual methods.
  • The proposed model outperformed existing models by about 2% in accuracy.

Abstract

Viruses are a significant threat to human life, as demonstrated by the global COVID-19 pandemic and the Ebola outbreak. Diseases such as smallpox, AIDS, hepatitis, liver cancer, and cervical cancer often caused by the Human Papillomavirus can lead to fatal outcomes. Over the years, extensive research has focused on developing vaccines and antiviral drugs, which have successfully contained and, in some cases, eradicated viral infections. Recently, computational techniques, particularly machine learning algorithms, have made notable progress in identifying potential Antiviral Peptides (AVPs), thereby accelerating experimental validation for therapeutic applications. This study introduces machine learning techniques and a Light Gradient Boosting Machine (LGBM) model combined with three feature-encoding methods to predict whether a protein sequence contains effective antiviral peptides. Eight machine learning algorithms were evaluated using both single-feature and combined feature encodings. Among them, the lightweight LGBM model trained on combined encoded features achieved the best performance, with an accuracy of 98%, precision of 97%, recall of 98%, F1-score of 98%, and an AUC of 1.00. Compared to existing models, the proposed approach achieved approximately 2% higher accuracy using individual encoding methods and about 3% higher accuracy with combined features. The reliability and effectiveness of the proposed model highlight its potential value for pharmaceutical development and academic research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hasan et al. (2026) studied this question.

synapsesocial.com/papers/696c774feb60fb80d139595dhttps://doi.org/10.1109/tcbbio.2026.3654071
Ask AI
Helpful
Bookmark
Share
View Full Paper