One of the first steps of any data analysis is to identify and document the extent of the missing data. Failure to do so can lead to errors. For example, most statistical procedures automatically exclude observations that are missing values for any variables being analyzed, regardless of whether the analyst is cognizant of these exclusions. There is no one-size-fits-all solution for handling missing data; the optimal strategy depends on the study design, the goals of the analysis, and the pattern of the missing data. This article reviews methods for dealing with missing data and offers general guidelines for choosing the correct approach for a given analysis. The potential impact of missing data depends largely on why the data are missing. The best-case scenario is when the reasons for the data being missing are completely random. For example, perhaps a distracted study researcher inadvertently forgot to measure height for a random subset of participants. Statisticians term these data “missing completely at random” (MCAR). Because no systematic differences exist between study participants with and without the measurement, MCAR data will not bias the results. In contrast, when data are missing systematically, improper handling can introduce bias. For example, if women who earn a high salary are more likely to skip a survey question about income than are men who earn a high salary, then ignoring the missing data will artificially inflate male salaries relative to female salaries. It turns out that if we have measured all the reasons related to the missing data (e.g., gender and high versus low salaries), we can still avoid bias by adequately accounting for these factors. Statisticians call these data “missing at random,” or MAR. The name is a bit confusing, but the idea is that the data are missing randomly from within subgroups of participants that we can define (e.g., women who earn low salaries and men who earn high salaries). In contrast, if we have not measured all the reasons related to missing data, there is nothing we can do to completely avoid bias. Statisticians call these data “missing not at random,” or MNAR. No strategy for handling missing data works perfectly in the case of MNAR data. In reality, we can't know whether our data are MCAR, MAR, or MNAR. In fact, most real datasets likely contain a mix of all 3 types. However, it is useful for statisticians to consider all 3 mechanisms when devising guidelines for how to best handle missing data. It used to be common practice to fill in missing values with arbitrary numbers such as “99” or “88” or “–99” during data entry. I advise against this practice, because it can lead to errors. For example, imagine that a dataset contains the variable body mass index (BMI), and missing values have been entered as 99s. It would be too easy to unknowingly include these values when calculating statistics on BMI, such as the average BMI. Obviously, averaging in values of 99 will distort the mean (but maybe not enough to make the mistake obvious!). A better convention is to simply leave missing values as blank cells. If it is important to record the reasons for the missing data, these reasons can be stored in a second variable. The best strategy for dealing with missing data is to avoid it altogether through careful data collection and follow-up, as well as by resolving missing data after the fact (for example, by locating missing forms or recontacting study participants). However, it is usually impossible to avoid missing data altogether, and thus statistical approaches for handling missing data are needed. Because missing data are incredibly complex, statisticians cannot devise a universal set of guidelines that works for all cases. Rather, they run simulations to predict the optimal approach for a multitude of different theoretical scenarios. The optimal approach may differ depending on whether (1) the study is cross-sectional or longitudinal; (2) the study is observational or randomized; (3) the missing values are in outcome or predictor variables; (4) the affected variables are of primary or secondary importance to the analysis; (5) the number of missing values is small or large; and (6) the data are MCAR, MAR, or MNAR. See Table 1 for a summary of recommended methods by type of study, type of variable, and extent of missing data. The simplest and most widely used strategy for handling missing data is to drop incomplete observations from the analysis altogether. Researchers sometimes perform complete-case analyses inadvertently, because most statistical procedures automatically omit incomplete observations. Imagine that a researcher fits a 10-variable regression model on a dataset of 100 people; if each of these 10 variables is missing values for just 2 different people, 20% of the observations will be excluded whether or not the analyst is aware of it. Complete-case analysis will not bias the results if the data are MCAR, because the analyzed sample is a random subset of the complete sample. Complete-case analysis will also produce unbiased results for MAR data if one adjusts the analysis for all the covariates that predict missing data 1. However, because of the loss of sample size, complete-case analysis reduces statistical power and precision. In general, it is wasteful to drop observations that are only missing values on a handful of variables. Imputation can rescue these observations (see the next section). One exception is that, for observational studies that aim to test a particular hypothesis (e.g., “Is bone mineral density related to stress fractures in runners?”), it seems prudent to exclude observations that have zero data on either the primary predictor or the primary outcome variable. Deleting an observation may also be warranted if it is missing almost all data points. However, for randomized trials, all randomized participants should be included in the final analysis regardless of the extent of their missing data (at least for intention-to-treat analyses). In addition to dropping incomplete observations, researchers should consider dropping variables with substantial missing data (eg, missing for >10% of the sample) if they are of secondary importance to the analysis. In this case, it may not be worth the effort to impute the missing values. Single imputation methods replace missing values with a plausible guess. For example, if a female participant is missing data on height, we could assign her the mean height for women in the sample. Or, better still, we could predict her height more precisely using a multiple linear regression that incorporates other available attributes, such as weight. In this approach, we would fit a linear regression model to the sample, with height as the outcome and other measured variables as predictors; then we would plug in the woman's values for these other variables to predict her height. Single imputation methods are straightforward to apply. They yield unbiased results for MCAR data, and—if one uses all the variables related to missing data to impute the missing values—for MAR data 1. However, single imputation methods have a critical problem: They reduce the standard deviation of the imputed variable, because they ignore the uncertainty in our guesses. This is easy to see with a simple example: If I have several missing data points for height and I assign each one the exact same value (the mean for height), this approach will clearly reduce the overall variability in heights. In turn, the artificially low standard deviations will lead to artificially low standard errors and P values. For variables that are missing a small number of values (eg., <2% of observations), the impact of single imputation on the standard deviation will be negligible. For these cases, I advocate the use of single imputation because the extra effort involved in multiple imputation yields little benefit. However, for variables with a greater number of missing values, multiple imputation is preferred. Multiple imputation is similar to single imputation but properly accounts for the uncertainty in the imputed values. Rather than generate a single value for each missing data point, the researcher generates multiple values (typically 5 or 10). A linear regression model for predicting height actually specifies a distribution of heights. In single imputation, we default to the mean of this distribution; for multiple imputation, we instead generate random instances from the distribution. The multiple imputed values are stored in multiple datasets. Final analyses are performed on each dataset separately, and then the results are combined into single estimates of effect. Although this process sounds arduous, multiple imputation is surprisingly easy to implement. Modern statistical packages have built-in algorithms that do most of the work for you. The biggest challenge is deciding which variables to include in the imputation model (the regression model used to predict the missing values). In general, statisticians recommend being inclusive—that is, including any variables related to missing data, correlated with the imputed variable, or used in final statistical models (including the outcome variable when imputing a predictor). If the variables selected for the imputation model themselves contain missing data, this can increase complexity. Simulations show that multiple imputation along with maximum likelihood methods (described later) are optimal for most common missing data scenarios 1-4. Unfortunately, multiple imputation is vastly underused. One review found that just 1% of 1105 articles published in top epidemiology journals reported using multiple imputation 5. An alternative to imputation is to use statistical algorithms that incorporate rather than exclude incomplete observations. The full-information maximum likelihood method uses whatever data are available for a given observation. It performs as well as multiple imputation in simulations, but it is simpler—there is no need to worry about specifying an imputation model or dealing with multiple layers of missingness. Maximum likelihood methods for missing data lag behind multiple imputation methods in popularity, likely because they are not as widely integrated in modern statistical packages, but this situation is rapidly changing. For longitudinal data with missing outcomes, many readily available statistical procedures also handle missing data adeptly. For example, survival analysis techniques such as Kaplan-Meier and Cox regression allow subjects to contribute information for however long they are followed up. Similarly, generalized estimating equation and mixed models only drop participants from the model for the specific time periods they are missing, not for all time points. Simulations show that generalized estimating equation and mixed models perform as well as multiple imputation for analyzing longitudinal data with missing outcomes 3. Missing data represent a vexing problem in medical studies. The biggest mistakes arise when researchers fail to adequately identify and deal with missing data. There is no universal “best” strategy for handling missing data, and no approach can avoid bias if systematic reasons for the missing data exist but have not been measured (so-called “MNAR” data). Deleting observations or variables may be warranted in some cases. Single imputation is sufficient when the missing data are sufficiently sparse. For longitudinal data with missing outcomes, many readily available statistical models handle the data without the need for imputation. Multiple imputation and maximum likelihood methods perform optimally in most other scenarios. Modern statistical packages have made it easier than ever to use these methods, but they remain woefully underused. The difference between single and multiple imputation is that single imputation yields one imputed value per missing observation, whereas multiple imputation yields multiple values. The sample mean for height is 67.4 inches and the sample standard deviation is 4.6 inches.
No takes yet. Share an insight, caveat, or question.
Kristin L. Sainani (2015) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: