The convergence rate and final performance of common deep learning models significantly benefited from heuristics such as learning rate schedules, distillation, skip connections, and normalization layers. In the of theoretical underpinnings, controlled experiments aimed at these strategies can aid our understanding of deep learning and the training dynamics. Existing approaches for empirical rely on tools of linear interpolation and visualizations with reduction, each with their limitations. Instead, we revisit such of heuristics through the lens of recently proposed methods for loss and representation analysis, viz., mode connectivity and canonical analysis (CCA), and hypothesize reasons for the success of the. In particular, we explore knowledge distillation and learning rate of (cosine) restarts and warmup using mode connectivity and CCA. Our analysis suggests that: (a) the reasons often quoted for the success cosine annealing are not evidenced in practice; (b) that the effect of rate warmup is to prevent the deeper layers from creating training; and (c) that the latent knowledge shared by the teacher is disbursed to the deeper layers.
No takes yet. Share an insight, caveat, or question.
Gotmare et al. (2018) studied this question.