The advancement of long-read RNA sequencing technologies leads to a bright future for transcriptome analysis. Reconstructing alternative isoforms from long-read RNA data is crucial in transcriptomic studies and poses computational challenges. In this research, we study the heterogeneity of current long reads RNA-seq data, and categorize the sequencing data into two types according to the fraction of the alternative splicing events appeared in the extracted reliably supported sub-path of the splicing graph. Leveraging the heterogeneity, we develop TransGram, an algorithm to reconstruct transcriptome from long RNA-seq reads, which employs the traditional assemble strategy to reconstruct the transcript-representing paths within the splicing graph, and incorporates a machine-learning-based model trained on splice graph-derived features to eliminate false positive transcripts. TransGram supports both annotation-free and annotation-guided modes. Comprehensive evaluations across multiple datasets and conditions demonstrate that TransGram consistently outperforms widely used tools StringTie2, IsoQuant, Bambu, and ESPRESSO, in terms of precision and recall. Long-read RNA sequencing datasets exhibit heterogeneity that affects transcript reconstruction. Here, the authors characterize this heterogeneity and develop TransGram, which adapts reconstruction to different data types to improve accuracy.
No takes yet. Share an insight, caveat, or question.
Ren et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: