This article considers replicability of the performance of predictors across studies. We suggest a general approach to investigating this issue, based on ensembles of prediction models trained on different studies. We quantify how the common practice of training on a single study accounts in part for the observed challenges in replicability of prediction performance. We also investigate whether ensembles of predictors trained on multiple studies can be combined, using unique criteria, to design robust ensemble learners trained upfront to incorporate replicability into different contexts and populations.
No takes yet. Share an insight, caveat, or question.
Patil et al. (2018) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: