MOTIVATION: Hashing has been widely used for indexing, querying and rapid similarity search in many bioinformatics applications, including sequence alignment, genome and transcriptome assembly, k-mer counting and error correction. Hence, expediting hashing operations would have a substantial impact in the field, making bioinformatics applications faster and more efficient. RESULTS: We present ntHash, a hashing algorithm tuned for processing DNA/RNA sequences. It performs the best when calculating hash values for adjacent k-mers in an input sequence, operating an order of magnitude faster than the best performing alternatives in typical use cases. AVAILABILITY AND IMPLEMENTATION: ntHash is available online at http://www.bcgsc.ca/platform/bioinfo/software/nthash and is free for academic use. CONTACTS: hmohamadi@bcgsc.ca or ibirol@bcgsc.caSupplementary information: Supplementary data are available at Bioinformatics online.
No takes yet. Share an insight, caveat, or question.
Mohamadi et al. (2016) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: