PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 21, 2026Scientific Reports4 citationsOpen Access

Deep learning-based identification of N6-methyladenine sites via hybrid feature fusion and SHAP-driven feature selection

MAMai AlzamelIUIslam UddinSKSalman Khan

Key Points

  • This research aims to develop a deep learning framework for accurately identifying N6-methyladenine (6 mA) sites in genomic data.
  • Implemented a deep neural network (DNN) that integrates hybrid feature extraction for 6 mA site identification.
  • Applied SHAP for feature importance analysis to select relevant features for classification.
  • Evaluated model performance using 5-fold cross-validation on datasets from F. vesca and R. chinensis.
  • Achieved training accuracies of 97.95% for F. vesca and 96.15% for R. chinensis.
  • Independent test dataset accuracies were 97.10% (F. vesca) and 95.32% (R. chinensis).
  • Demonstrated strong generalisation capability across diverse genomic contexts.

Abstract

N6-methyladenine (6 mA) is a critical epigenetic modification involved in gene regulation, genome stability, and cellular adaptation. Accurate computational identification of 6 mA sites is essential for elucidating epigenetic mechanisms and advancing disease-related research. However, existing methods are often constrained by limited feature representations and poor interpretability, hindering predictive performance and generalisation across diverse genomic contexts. To address these challenges, we propose a novel deep neural network (DNN)-based framework that integrates optimal hybrid features for robust and interpretable 6 mA site identification. The proposed framework employs a comprehensive multi-feature extraction strategy to capture complex sequence-level patterns in DNA, which are subsequently fused into a unified hybrid representation. To enhance model efficiency and interpretability, SHapley Additive exPlanations (SHAP) are applied for feature importance analysis, enabling the selection of the most discriminative features for downstream classification. The optimised feature set is then fed into a DNN classifier for accurate 6 mA site prediction. Evaluated using 5-fold cross-validation, the proposed model achieved training accuracies of 97.95% and 96.15% on the F. vesca and R. chinensis datasets, respectively. On independent test datasets, the model achieved accuracies of 97.10% (F. vesca) and 95.32% (R. chinensis), demonstrating strong generalisation. These results establish the proposed framework as an accurate, reliable, and interpretable computational tool for genome-wide identification of 6 mA sites, with broad applicability to epigenetic research and beyond.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Alzamel et al. (2026) studied this question.

synapsesocial.com/papers/6a0ea0f7be05d6e3efb5f4a0https://doi.org/10.1038/s41598-026-50667-z
Ask AI
Helpful
Bookmark
Share
View Full Paper