PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 2, 2018Biometrical Journal1,719 citationsOpen Access

Variable selection – A review and recommendations for the practicing statistician

View Full Paper
GHGeorg HeinzeUniversity of ViennaCWChristine WallischMedical University of ViennaDDDaniela DunklerMedical University of Vienna

Key Points

  • The aim is to review variable selection methods and their implications for unbiased effect estimates in statistical modeling.
  • Overview of variable selection methods including significance, information criteria, and penalized likelihood.
  • Examination of the impact of variable selection on regression coefficients and confidence intervals.
  • Recommendations for using variable selection in low-dimensional modeling and stability investigations.
  • Variable selection can compromise model stability and unbiasedness of effect estimates.
  • Proposed resampling quantities to be reported with automated selection algorithms enhance model reliability.
  • Emphasis on pragmatic guidance for statisticians to apply variable selection effectively.

Abstract

Statistical models support medical research by facilitating individualized outcome prognostication conditional on independent variables or by estimating effects of risk factors adjusted for covariates. Theory of statistical models is well-established if the set of independent variables to consider is fixed and small. Hence, we can assume that effect estimates are unbiased and the usual methods for confidence interval estimation are valid. In routine work, however, it is not known a priori which covariates should be included in a model, and often we are confronted with the number of candidate variables in the range 10-30. This number is often too large to be considered in a statistical model. We provide an overview of various available variable selection methods that are based on significance or information criteria, penalized likelihood, the change-in-estimate criterion, background knowledge, or combinations thereof. These methods were usually developed in the context of a linear regression model and then transferred to more generalized linear models or models for censored survival data. Variable selection, in particular if used in explanatory modeling where effect estimates are of central interest, can compromise stability of a final model, unbiasedness of regression coefficients, and validity of p-values or confidence intervals. Therefore, we give pragmatic recommendations for the practicing statistician on application of variable selection methods in general (low-dimensional) modeling problems and on performing stability investigations and inference. We also propose some quantities based on resampling the entire variable selection process to be routinely reported by software packages offering automated variable selection algorithms.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Heinze et al. (2018) studied this question.

synapsesocial.com/papers/69d720c73c36f67a08563787https://doi.org/10.1002/bimj.201700067
Ask AI
Helpful
Bookmark
Share
View Full Paper