Key points are not available for this paper at this time.
Identifying informative predictors in a high dimensional regression model is critical step for association analysis and predictive modeling. Signal in the high dimensional setting often fails due to the limited sample. One approach to improve power is through meta-analyzing multiple studies the same scientific question. However, integrative analysis of high data from multiple studies is challenging in the presence of study heterogeneity. The challenge is even more pronounced with data sharing constraints under which only summary data but not level data can be shared across different sites. In this paper, we a novel data shielding integrative large-scale testing (DSILT) approach signal detection by allowing between study heterogeneity and not requiring of individual level data. Assuming the underlying high dimensional models of the data differ across studies yet share similar support, DSILT approach incorporates proper integrative estimation and debiasing to construct test statistics for the overall effects of specific. We also develop a multiple testing procedure to identify effects while controlling for false discovery rate (FDR) and false proportion (FDP). Theoretical comparisons of the DSILT procedure with ideal individual--level meta--analysis (ILMA) approach and other inference methods are investigated. Simulation studies demonstrate the DSILT procedure performs well in both false discovery control and power. The proposed method is applied to a real example on detecting effect of the genetic variants for statins and obesity on the risk Type 2 Diabetes.
Liu et al. (Thu,) studied this question.