Computational tools are widely used for interpreting variants detected in sequencing projects.The choice of these tools is critical for reliable variant impact interpretation for precision medicine and should be based on systematic performance assessment.The performance of the methods varies widely in different performance assessments, for example due to the contents and sizes of test datasets.To address this issue, we obtained 63,160 common amino acid substitutions (allele frequency �1% and <25%) from the Exome Aggregation Consortium (ExAC) database, which contains variants from 60,706 genomes or exomes.We evaluated the specificity, the capability to detect benign variants, for 10 variant interpretation tools.In addition to overall specificity of the tools, we tested their performance for variants in six geographical populations.PON-P2 had the best performance (95.5%) followed by FATHMM (86.4%) and VEST (83.5%).While these tools had excellent performance, the poorest method predicted more than one third of the benign variants to be disease-causing.The results allow choosing reliable methods for benign variant interpretation, for both research and clinical purposes, as well as provide a benchmark for method developers. Author summaryIn precision/personalized medicine of many conditions it is essential to investigate individual's genome.Interpretation of the observed variation (mutation) sets is feasible only with computational approaches.We assessed the performance of variant pathogenicity/ tolerance prediction programs on benign variants.Variants were obtained from highquality ExAC database and selected to have minor allele frequency between 1 and 25%.We obtained 63,160 such cases and investigated 10 widely used predictors.Specificities of the methods showed large differences, from 64 to 96%, thus users of these methods have to be careful when choosing the one(s) they will use.We investigated further the performances on different populations, allele frequencies, separately for males and females, chromosome wise and for population unique and non-unique variants.The ranking of the tools remained the same in all these scenarios, i.e. the best methods were the best irrespective on how the data was filtered and grouped.This is to our knowledge the first large scale evaluation of method performance on benign variants.
No takes yet. Share an insight, caveat, or question.
Niroula et al. (2019) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: