PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 30, 20250 citationsOpen Access

The Hidden Width of Deep ResNets: Tight Error Bounds and Phase Diagrams

View Full Paper
LCLénaïc Chizat

Key Points

  • The training dynamics of ResNets converge to Neural Mean ODEs, showing a clear dependency on network depth and width.
  • An error bound of O_D(1/L + α/√(LM)) is established for the model's output with respect to its limit after specific gradient steps.
  • Complete feature learning is exhibited in a defined regime, highlighting that only certain residual scales enable this rate of learning.
  • Empirical verification reinforces the theoretical findings regarding error bounds for two-layer perceptron blocks in ResNets.

Abstract

We study the gradient-based training of large-depth residual networks (ResNets) from standard random initializations. We show that with a diverging depth L, a fixed embedding dimension D, and an arbitrary hidden width M, the training dynamics converges to a Neural Mean ODE training dynamics. Remarkably, the limit is independent of the scaling of M, covering practical cases of, say, Transformers, where M (the number of hidden units or attention heads per layer) is typically of the order of D. For a residual scale ΘD (αLM), we obtain the error bound OD (1L+ αLM) between the model's output and its limit after a fixed number gradient of steps, and we verify empirically that this rate is tight. When α=Θ (1), the limit exhibits complete feature learning, i. e. the Mean ODE is genuinely non-linearly parameterized. In contrast, we show that α yields a ODE regime where the Mean ODE is linearly parameterized. We then focus on the particular case of ResNets with two-layer perceptron blocks, for which we study how these scalings depend on the embedding dimension D. We show that for this model, the only residual scale that leads to complete feature learning is Θ (DLM). In this regime, we prove the error bound O (1L+ DLM) between the ResNet and its limit after a fixed number of gradient steps, which is also empirically tight. Our convergence results rely on a novel mathematical perspective on ResNets: (i) due to the randomness of the initialization, the forward and backward pass through the ResNet behave as the stochastic approximation of certain mean ODEs, and (ii) by propagation of chaos (that is, asymptotic independence of the units) this behavior is preserved through the training dynamics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lénaïc Chizat (2025) studied this question.

synapsesocial.com/papers/68dc1e358a7d58c25ebb16afhttps://doi.org/10.48550/arxiv.2509.10167
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Generalization of Scaled Deep ResNets in the Mean-Field Regime2024
  2. 2Understanding the training of infinitely deep and wide ResNets with Conditional Optimal Transport2024
  3. 3Exact solutions to the nonlinear dynamics of learning in deep linear neural networks2013 · 1,010 citations
  4. 4Infinite‐width limit of deep linear neural networks2024 · 7 citations
  5. 5Hamiltonian Mechanics of Feature Learning: Bottleneck Structure in Leaky ResNets2024