PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 9, 20242 citationsOpen Access

Unifying Low Dimensional Observations in Deep Learning Through the Deep Linear Unconstrained Feature Model

View Full Paper
CGConnall GarrodJKJonathan P. Keating

Key Points

  • Theoretical analysis reveals unification of low-dimensional structures across various neural network components.
  • Observed occurrences include low-dimensional behavior in gradients and Hessian spectra during convergence.
  • The analysis is based on a generalized unconstrained feature model, addressing Neural Collapse and its multi-layer version conceptually at optimal points in training periods, like global optima, found in various settings and architectures. The empirical support from both models indicates validation of these theories.

Abstract

Modern deep neural networks have achieved high performance across various tasks. Recently, researchers have noted occurrences of low-dimensional structure in the weights, Hessian's, gradients, and feature vectors of these networks, spanning different datasets and architectures when trained to convergence. In this analysis, we theoretically demonstrate these observations arising, and show how they can be unified within a generalized unconstrained feature model that can be considered analytically. Specifically, we consider a previously described structure called Neural Collapse, and its multi-layer counterpart, Deep Neural Collapse, which emerges when the network approaches global optima. This phenomenon explains the other observed low-dimensional behaviours on a layer-wise level, such as the bulk and outlier structure seen in Hessian spectra, and the alignment of gradient descent with the outlier eigenspace of the Hessian. Empirical results in both the deep linear unconstrained feature model and its non-linear equivalent support these predicted observations.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Garrod et al. (2024) studied this question.

synapsesocial.com/papers/68e6febab6db643587678e92https://doi.org/10.48550/arxiv.2404.06106
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?2024
  2. 2Geometric Analysis of Unconstrained Feature Models with $d=K$2024
  3. 3Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers2025
  4. 4Low-Rank Learning by Design: the Role of Network Architecture and Activation Linearity in Gradient Rank Collapse2024
  5. 5An Analytical Characterization of Sloppiness in Neural Networks: Insights from Linear Models2025