Benchmarking study demonstrates reduced false-positive strain identification across complex metagenomic datasets, indicating improved accuracy for pathogen surveillance.
Key Points
To resolve ambiguous read assignments among closely related microbial genomes and reduce false-positive strain detections in metagenomic sequencing data.
Developed StrainRefine, a post-mapping algorithm that constructs binary read-support profiles for candidate reference genomes to quantify profile similarity.
Clustered genomes based on mapping profiles, filtered out weakly supported candidates, and reassigned reads to representative references without relying on prior sample composition assumptions.
Evaluated read-level accuracy, precision-recall balance, and abundance profile concordance against existing mapping-based approaches across large-scale and single-species metagenomic datasets.
Substantially decreased false-positive strain identifications while maintaining recall and improving agreement between predicted and true abundance profiles.
Attained the highest read-level classification accuracy on the most complex evaluated benchmark dataset compared with existing tools.
Demonstrated performance on par with species-specific tools without requiring curated species-specific reference databases or prior compositional assumptions.