Key points are not available for this paper at this time.
Several gene-based tests, such as the sequence kernel association test, have been developed to assess associations between rare single nucleotide variants (SNVs) and disease traits. However, these aggregate methods do not distinguish potentially causal variants from null variants within associated regions. To address this limitation, we propose gvClust, a clustering approach that classifies rare variants into null and signal groups using a Gaussian mixture model applied to variant-level summary statistics from multiple-variant models. Signal variants are further partitioned into risk and protective subgroups according to their effect direction and magnitude. We evaluated gvClust in simulation studies using the adjusted Rand index (ARI), mean squared error (MSE), and accuracy of cluster number selection under different sample sizes, effect configurations, outcome types, and linkage disequilibrium (LD) structures. In simulations, gvClust showed improved performance with increasing sample size, achieved high accuracy in determining the number of clusters for continuous traits at large sample sizes, and outperformed both k-means clustering and initialization-only clustering. We then applied gvClust to rare variants in six genes associated with blood pressure traits from a large genome-wide association study and meta-analysis. In the real-data application, gvClust identified distinct null, risk, and protective clusters. These results suggest that gvClust provides a practical framework for classifying rare variants within associated regions and may help improve the biological interpretation of rare variant signals.
Sun et al. (Sun,) studied this question.