Recent developments in applications of artificial neural networks with over n = 10 14 parameters make it extremely important to study the large n behavior of such networks. Most works studying wide neural networks have focused on the infinite width n → + ∞ limit of such networks and have shown that, at initialization, they correspond to Gaussian processes [12] , [23] . In this work we will study their behavior for large, but finite n . Our main contributions are the following: • The computation of the corrections to Gaussianity in terms of an asymptotic series in n − 1 2 . The coefficients in this expansion are determined by the statistics of parameter initialization and by the activation function. • The quantitative control, in terms of the width n , of the evolution of the outputs of (finite) networks, during training, by computing deviations from the limiting infinite width case (in which the network evolves through a linear flow). This as been the subject of several previous works, notably [2] , [10] , [16] , [21] , [25] ; our results improve previous estimates and along the way also provide sharper decay rates for the (finite width) NTK in terms of n , valid during the entire training procedure. • Estimating how the deviations from Gaussianity evolve with training in terms of n . In particular, using a certain metric in the space of measures we find that, along training, the resulting measure is within n − 1 2 ( log n ) + of the time dependent Gaussian process corresponding to the infinite width network (which is explicitly given by precomposing the initial Gaussian process with the linear flow corresponding to training in the infinite width limit).
No takes yet. Share an insight, caveat, or question.
Carvalho et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: