This paper proposes an analysis of case-control data under a double-sampling scheme, when covariates are missing or measured with error at the first stage of sampling and are validated at the second stage in a subsample. The method combines risk information from both samples. It is derived under the assumptions that (i) the prospective disease incidence model is of logistic form, (ii) the proxy or partial information takes on finitely many values, and (iii) the error is nondifferential. The method of Prentice & Pyke (1979) is extended to a two-sample design to derive an estimating equation for the odds-ratio parameters. It provides an alternative estimator to that given by Breslow & Cain (1988). Consistency and asymptotic normality of the estimates are derived and a variance formula is presented. Parameters can be estimated by use of standard packages for quantal-response data. The method can easily be extended to the analysis of stratified designs with large strata.
No takes yet. Share an insight, caveat, or question.
Schill et al. (1993) studied this question.