We prove that for an L-layer fully-connected linear neural network, if the width of every hidden layer is Ω(L · r · dₒᵤₜ · κ³ ), where r and $κ$ are the rank and the condition number of the input data, and dₒᵤₜ is the output dimension, then gradient descent with Gaussian random initialization converges to a global minimum at a linear rate. The number of iterations to find an $ε$-suboptimal solution is O(κlog(1ε)). Our polynomial upper bound on the total running time for wide deep linear networks and the exp(Ω(L)) lower bound for narrow deep linear neural networks [Shamir, 2018] together demonstrate that wide layers are necessary for optimizing deep models.
No takes yet. Share an insight, caveat, or question.
Du et al. (2019) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: