In this paper, we propose a model-free feature screening framework tailored for high-dimensional and heterogeneous datasets, based on a novel distributionally robust dependence measure termed Copula Divergence. The proposed screening method, named CD-Screen, addresses critical limitations of existing feature screening methods, such as restrictive modeling assumptions and sensitivity to heterogeneous feature distributions. CD-Screen ranks features according to their Copula Divergence without relying on a specific regression model or distributional assumptions. Additionally, we introduce CD-FDR, a data-driven procedure to control false discoveries, ensuring accurate and efficient feature selection. Theoretical analyses establish the sure screening and rank consistency properties of CD-Screen, along with asymptotic control of the false discovery rate by CD-FDR. Extensive simulation studies demonstrate the superior performance of our methods compared to traditional screening approaches across diverse scenarios. Furthermore, a real data analysis of the relationship between stock returns and inflation in the United States illustrates the practical use of our method and provides descriptive evidence on sector-specific responses to economic changes.
Cheng et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: