Authors
Background: It is still an open research area to theoretically understand why Neural Networks (DNNs)---equipped with many more parameters than training and trained by (stochastic) gradient-based methods---often achieve low generalization error. Contribution: We study DNN training by analysis. Our theoretical framework explains: i) DNN with (stochastic)-based methods often endows low-frequency components of the target with a higher priority during the training; ii) Small initialization to good generalization ability of DNN while preserving the DNN's ability fit any function. These results are further confirmed by experiments of DNNs the following datasets, that is, natural images, one-dimensional and MNIST dataset.
No takes yet. Share an insight, caveat, or question.
Zhi‐Qin John Xu (2018) studied this question.