Key points are not available for this paper at this time.
Abstract Metagenomic sequencing is transforming diverse areas of health and biological sciences, including pathogen surveillance, clinical diagnostics, and microbiome research. However, the inherent complexity of metagenomic data limits most computational tools to species-level classification and abundance estimation, overlooking within-species genetic diversity that drives key phenotypes. We present metaWEPP, a novel computational pipeline that achieves near-haplotype resolution in metagenomic analysis for species with adequate representation in reference genome biobanks and having sufficient sequencing depth and genome coverage. Specifically, metaWEPP assigns sequencing reads to species using standard taxonomic classifiers, phylogenetically places them onto species-specific mutation-annotated trees of publicly available sequences, and selects the haplotypes that best explain the sample. It also reports unaccounted alleles indicative of novel variants and provides an interactive dashboard for read-level visualization. Applied to diverse metagenomic and mixed-genome samples from prior studies, metaWEPP produced concordant species-level results, while revealing finer lineage- and haplotype-level insights not captured by existing tools. On various clinical samples, metaWEPP identified infecting pathogens and additionally provided credible lineage- and haplotype-level information that can support clinical decision-making. On wastewater samples, metaWEPP uncovered previously undetected haplotype clusters of epidemiological relevance. These findings demonstrate metaWEPP’s ability to advance various clinical, epidemiological, and research applications with deeper, actionable insights.
Gangwar et al. (Sat,) studied this question.