Numerous empirical evidence has corroborated that the noise plays a crucial in effective and efficient training of neural networks. The theory behind,, is still largely unknown. This paper studies this fundamental problem training a simple two-layer convolutional neural network model. training such a network requires solving a nonconvex optimization with a spurious local optimum and a global optimum, we prove that gradient descent and perturbed mini-batch stochastic gradient in conjunction with noise annealing is guaranteed to converge to a optimum in polynomial time with arbitrary initialization. This implies the noise enables the algorithm to efficiently escape from the spurious optimum. Numerical experiments are provided to support our theory.
No takes yet. Share an insight, caveat, or question.
Zhou et al. (2019) studied this question.