In a number of special cases, it has been shown that categorization of a continuous explanatory variable for use in a regression model can lead to efficiency losses. Nevertheless, categorization remains a popular means of dealing with continuous variables, particularly in medical statistics. Two topics of current interest in medical statistics are subgroup effects in clinical trials and gene–environment interactions in genetic studies. In a regression model setting, these topics involve the examination of interaction effects, often between a binary explanatory variable and a continuous explanatory variable. In this article the efficiency losses associated with dichotomization of a continuous explanatory variable for the testing of such interaction effects are calculated and compared with those associated with the testing of main effects. It is shown that considerable additional efficiency loss can arise because of dichotomization in both a main effect and an interaction. The theoretical development is done in the context of exponential family models, thus also generalizing earlier results on efficiency loss for main effects. Some indication is also given of the losses in efficiency associated with more detailed categorization. The further impact of censoring in time-to-event models is investigated and shown to not alter the qualitative conclusions. The practical importance of the findings is illustrated through the analysis of data from a clinical trial in patients with prostate cancer.
No takes yet. Share an insight, caveat, or question.
Farewell et al. (2004) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: