Abstract Automated implementations of the ACMG/AMP variant classification guidelines are increasingly used to support clinical genomics, yet systematic comparisons with expert human curation remain limited. In this study, we benchmarked several widely used automated and AI‑assisted tools, including Franklin, VarSome, MobiDetails, GeneBe, InterVar, and VarChat, against dual independent curator assessments across diverse variant types. We quantified criterion‑level agreement, evidence weighting behavior, and classification concordance, with a particular focus on calibration around key ACMG criteria. Our analyses revealed that discrepancies between tools and curators concentrated around evidence‑strength calibration and near‑boundary categories (LP↔P, LP↔VUS). Loss‑of‑function variants showed the highest concordance, reflecting the maturity of PVS1‑based decision trees, whereas missense and splicing variants exhibited wider variability driven by differences in PM1 hotspot definitions, PP3/BP4 predictor thresholds, and access to case‑level and segregation evidence. Tool‑specific patterns were evident: Franklin and VarSome demonstrated high concordance but a slight pathogenic-leaning bias; VarChat showed near‑neutral calibration; GeneBe yielded higher and more variable evidence totals; and InterVar applied more conservative, lower‑weight scoring. Overall, our findings indicate that automated tools perform reliably for structured, data‑rich evidence but benefit from expert adjudication for context‑dependent criteria. This is especially relevant as laboratories prepare for the forthcoming ACMG v4 framework while still operating under established ACMG and ACGS recommendations, creating a transitional period in which robust and transparent workflows remain essential.
Slapnik et al. (Wed,) studied this question.