PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 7, 2026Sensors0 citationsOpen Access

Ensemble Deep Learning Models on Raw DNA Sequences for Viral Genome Identification in Human Samples

View Full Paper
MNMarco De NatSBSimone BoscoloSGSonia Pilar Gallo

Key Points

  • This research aims to develop a deep learning framework for identifying viral genomes in complex human samples.
  • Created an ensemble of deep convolutional neural networks for viral identification.
  • Processed high-throughput biological sensor data.
  • Integrated different network architectures to capture local and global genomic features.
  • Evaluated model performance using AUROC metrics.
  • Achieved an AUROC of 0.939 on 300 bp viral contigs.
  • Outperformed existing models like transformer-based methods and ViraMiner.
  • Maintained predictive power with a 10% random nucleotide substitution.
  • Generalized effectively to unseen viral families.

Abstract

Detecting highly divergent or previously unknown viruses is a critical bottleneck in clinical diagnostics and pathogen surveillance. While alignment-based methods often fail to classify sequences lacking homology to known references, deep learning offers a powerful alternative for signal extraction from ‘viral dark matter.’ In this work, we present a high-performance ensemble of deep convolutional neural networks specifically designed to identify viral contigs in complex human metagenomic datasets. Our framework processes sequences acquired from high-throughput biological sensors and integrates complementary architectures to capture both local motifs and global genomic signatures. The proposed ensemble achieves state-of-the-art performance, reaching an AUROC of 0.939 on 300 bp contigs and significantly outperforming existing models such as transformer-based approaches, ViraMiner, and DeepVirFinder. Crucially, our results demonstrate high robustness to data degradation, maintaining stable predictive power even with a 10% random nucleotide substitution rate, a common challenge in degraded clinical samples. Furthermore, the model generalizes to ‘unseen’ viral families not present during training, demonstrating its utility for emerging threat detection. To ensure full reproducibility and facilitate further research in clinical sensing, the complete code and datasets are publicly available on Github.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nat et al. (2026) studied this question.

synapsesocial.com/papers/69d49f44b33cc4c35a227b19https://doi.org/10.3390/s26072238
Ask AI
Helpful
Bookmark
Share
View Full Paper