Randomized trial compares new MMRDA method to traditional RA in evaluating imputation techniques, suggesting better assessment strategies.
In data science, Missing Data (MD) are handled by approaches from basic imputation ( e.g., mean) to more complex ( e.g., Machine Learning (ML)). Assessing these methods consists in amputating data then imputating back amputated values to compare with ground truth. The existing amputation methods are most often Random Amputation (RA) or parameterized approaches using Missing Data Mechanisms (MDM) involved. MD from real-world datasets can follow complex patterns and the specific MDM involved can be impossible to determine with certainty. Thus, existing data amputation methods are often difficult to properly apply on real-world datasets. In this article, we propose Missing Mechanisms Respectful Data Amputation (MMRDA) as a new method to generate synthetic MD and we study the impact of data amputation methods on the results of imputation techniques assessment. MMRDA has been compared to the RA method and assessed on open datasets. MMRDA significantly outperforms RA. Four well-known imputation techniques were used to compare imputation performance post-MMRDA with imputation performance post-RA. Differences were significant and the best imputation performance is not always achieved by the same imputer depending on the amputation technique used for a given dataset. This supports that MDM involved affect imputation and that the attribution of complex mechanisms is still to be explored. Therefore, we recommend all researchers to pay close attention to the amputation method used during imputation techniques assessment or comparison, especially on real-world datasets.
No takes yet. Share an insight, caveat, or question.
Pitteman et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: