PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 30, 2016Scientific Reports352 citationsOpen Access

DBG2OLC: Efficient Assembly of Large Genomes Using Long Erroneous Reads of the Third Generation Sequencing Technologies

CYChengxi YeCHC. HillSWShigang Wu

Key Points

  • To develop a fast and memory-efficient hybrid genome assembly algorithm (DBG2OLC) that integrates short accurate reads and long noisy reads for large genomes.
  • Derived a compact representation of third generation sequencing (3GS) long reads using preassembled next generation sequencing (NGS) contigs, converting a de Bruijn graph into an overlap graph.
  • Identified and corrected structural errors in long reads while bypassing base-level error correction prior to assembly.
  • Assembled and polished structurally verified 3GS reads using complementary NGS data on mammalian-sized genomic datasets.
  • Assembled mammalian-sized genomes orders of magnitude faster than existing assembly pipelines with low memory consumption.
  • Reduced sequencing coverage requirements for both sequencing technologies, cutting overall sequencing costs by approximately 50%.

Abstract

The highly anticipated transition from next generation sequencing (NGS) to third generation sequencing (3GS) has been difficult primarily due to high error rates and excessive sequencing cost. The high error rates make the assembly of long erroneous reads of large genomes challenging because existing software solutions are often overwhelmed by error correction tasks. Here we report a hybrid assembly approach that simultaneously utilizes NGS and 3GS data to address both issues. We gain advantages from three general and basic design principles: (i) Compact representation of the long reads leads to efficient alignments. (ii) Base-level errors can be skipped; structural errors need to be detected and corrected. (iii) Structurally correct 3GS reads are assembled and polished. In our implementation, preassembled NGS contigs are used to derive the compact representation of the long reads, motivating an algorithmic conversion from a de Bruijn graph to an overlap graph, the two major assembly paradigms. Moreover, since NGS and 3GS data can compensate for each other, our hybrid assembly approach reduces both of their sequencing requirements. Experiments show that our software is able to assemble mammalian-sized genomes orders of magnitude more quickly than existing methods without consuming a lot of memory, while saving about half of the sequencing cost.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ye et al. (2016) studied this question.

synapsesocial.com/papers/6a154405d64fa333899f7361https://doi.org/10.1038/srep31900
Ask AI
Helpful
Bookmark
Share
View Full Paper