Abstract DNA barcoding and metabarcoding have emerged as cost‐efficient, standardized methods for characterizing local biodiversity. Based on the sequencing of a small targeted gene fragment, it is theoretically possible to identify a wide diversity of taxa by comparing them with reference sequence databases. However, a key challenge for accurate taxonomic classification is the incompleteness of such databases, leading to most query sequences lacking species‐level matches. Where species‐level matches are missing, it may be possible to classify query sequences to a higher taxonomic rank, such as genus or family, based on the similarity of related reference taxa. The challenge then lies in confidently recognizing whether the sequence belongs to an unobserved (here, ‘novel’) taxon on a given rank. We evaluate the performance and utility of several methods for taxonomic classification. Methods were assessed based on the classification accuracy of both observed and novel taxa, accuracy of prediction confidence estimates and computational resource use. We focus on two widely studied cases: the COI barcode for arthropods, and the ITS barcode for fungi, with the latter representing an instance with substantially greater sequence length variation within classes. To benchmark the classification of novel taxa, we used curated datasets with partially distinct taxonomic distributions between the training and test sets. Novel taxa occurred at all evaluated taxonomic ranks, such as novel species in observed genera and novel genera in observed families. We further assessed the effect on performance when shifting from full‐length barcodes to shorter sequences as generated through metabarcoding in the test dataset. This study sheds light on the strengths and limitations of different classification algorithms across taxonomic groups and barcode characteristics. It demonstrates the supreme performance of phylogenetic placement methods (e.g., EPA‐ng) for classification of arthropod COI barcodes and composition‐based classifiers (e.g., SINTAX, RDP‐NBC, IDTAXA) for fungal ITS. Differences likely reflect barcode properties: COI is alignable and evolutionary constrained, favouring phylogenetic placement, whereas ITS is too variable for reliable alignment but rich in short‐motif (k‐mer) signal, favouring composition‐based classifiers. Across most algorithms, shorter sub‐regions performed comparably to full‐length barcodes, supporting the use of short reads for high‐throughput metabarcoding.
Orsholm et al. (Thu,) studied this question.