A main problem with reproducing machine learning publications is the variance of metric implementations across papers. A lack of standardization leads to different behavior in mechanisms such as checkpointing, learning rate schedulers or early stopping, that will influence the reported results. For example, a complex metric such as Frchet inception distance (FID) for synthetic image quality evaluation
No takes yet. Share an insight, caveat, or question.
Detlefsen et al. (2022) studied this question.