Background: Nanopore sequencing produces ionic current signals that are sensitive to chemical modifications in DNA and RNA molecules. However, accurate modification detection remains challenging due to limited labeled data and variability across experimental conditions. Methods: We present a scalable unsupervised framework for modification discovery that learns reference signal distributions from unmodified sequences using a CNN–Transformer variational autoencoder (VAE). The model is trained on large-scale data via streaming sampling and k-mer-aware soft balancing to ensure robust signal representation. At inference, candidate nucleotides are scored using the VAE reconstruction error, and read-level signals are aggregated to produce site-level modification evidence. Results: On controlled DNA oligonucleotide datasets, models trained on unmodified sequences achieve strong discrimination when evaluated on modified oligos. In contrast, performance decreases in cell line samples when models trained on unmodified whole-genome-amplified (WGA) DNA and in vitro-transcribed (IVT) RNA are evaluated on natively modified (5mC/m6A) data, reflecting the impacts of biological noise and heterogeneity. Despite reduced classification accuracy, site-level anomaly score profiles exhibit peak-like patterns that correspond to known modification-enriched regions. Conclusions: These findings demonstrate the feasibility of large-scale unsupervised reference modeling for de novo modification detection, while underscoring the challenges in translating models built from synthetic oligo datasets into robust genome-wide modification detection.
Zou et al. (Wed,) studied this question.