One of the fundamental assumptions behind many supervised machine‐learning algorithms is that training and test data follow the same probability distribution. However, this important assumption is often violated in practice, for example, because of an unavoidable sample selection bias or nonstationarity of the environment. Owing to violation of the assumption, standard machine‐learning methods suffer a significant estimation bias. In this article, we consider two scenarios of such distribution change—the covariate shift where input distributions differ and class‐balance change where class‐prior probabilities vary in classification—and review semi‐supervised adaptation techniques based on importance weighting . WIREs Comput Stat 2013, 5:465–477. doi: 10.1002/wics.1275 This article is categorized under: Statistical Learning and Exploratory Methods of the Data Sciences > Clustering and Classification Statistical Learning and Exploratory Methods of the Data Sciences > Pattern Recognition
No takes yet. Share an insight, caveat, or question.
Sugiyama et al. (2013) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: