Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Min, Hancheng, Zhu, Zhihui, Vidal, René
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918169231228928
author Min, Hancheng
Zhu, Zhihui
Vidal, René
author_facet Min, Hancheng
Zhu, Zhihui
Vidal, René
contents Among many mysteries behind the success of deep networks lies the exceptional discriminative power of their learned representations as manifested by the intriguing Neural Collapse (NC) phenomenon, where simple feature structures emerge at the last layer of a trained neural network. Prior works on the theoretical understandings of NC have focused on analyzing the optimization landscape of matrix-factorization-like problems by considering the last-layer features as unconstrained free optimization variables and showing that their global minima exhibit NC. In this paper, we show that gradient flow on a two-layer ReLU network for classifying orthogonally separable data provably exhibits NC, thereby advancing prior results in two ways: First, we relax the assumption of unconstrained features, showing the effect of data structure and nonlinear activations on NC characterizations. Second, we reveal the role of the implicit bias of the training dynamics in facilitating the emergence of NC.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21078
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
Min, Hancheng
Zhu, Zhihui
Vidal, René
Machine Learning
Optimization and Control
Among many mysteries behind the success of deep networks lies the exceptional discriminative power of their learned representations as manifested by the intriguing Neural Collapse (NC) phenomenon, where simple feature structures emerge at the last layer of a trained neural network. Prior works on the theoretical understandings of NC have focused on analyzing the optimization landscape of matrix-factorization-like problems by considering the last-layer features as unconstrained free optimization variables and showing that their global minima exhibit NC. In this paper, we show that gradient flow on a two-layer ReLU network for classifying orthogonally separable data provably exhibits NC, thereby advancing prior results in two ways: First, we relax the assumption of unconstrained features, showing the effect of data structure and nonlinear activations on NC characterizations. Second, we reveal the role of the implicit bias of the training dynamics in facilitating the emergence of NC.
title Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2510.21078