Optimizer choice matters for the emergence of Neural Collapse
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Jim, Cheng, Tin Sum, Masarczyk, Wojciech, Lucchi, Aurelien |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization
by: Masarczyk, Wojciech, et al.
Published: (2025)
by: Masarczyk, Wojciech, et al.
Published: (2025)
Characterizing Overfitting in Kernel Ridgeless Regression Through the Eigenspectrum
by: Cheng, Tin Sum, et al.
Published: (2024)
by: Cheng, Tin Sum, et al.
Published: (2024)
A Comprehensive Analysis on the Learning Curve in Kernel Ridge Regression
by: Cheng, Tin Sum, et al.
Published: (2024)
by: Cheng, Tin Sum, et al.
Published: (2024)
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters
by: Kratsios, Anastasis, et al.
Published: (2025)
by: Kratsios, Anastasis, et al.
Published: (2025)
Cubic regularized subspace Newton for non-convex optimization
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
by: Islamov, Rustem, et al.
Published: (2024)
by: Islamov, Rustem, et al.
Published: (2024)
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
Optimization Guarantees for Square-Root Natural-Gradient Variational Inference
by: Kumar, Navish, et al.
Published: (2025)
by: Kumar, Navish, et al.
Published: (2025)
Spectral Analysis of Molecular Kernels: When Richer Features Do Not Guarantee Better Generalization
by: Jamali, Asma, et al.
Published: (2025)
by: Jamali, Asma, et al.
Published: (2025)
A Theoretical Analysis of the Learning Dynamics under Class Imbalance
by: Francazi, Emanuele, et al.
Published: (2022)
by: Francazi, Emanuele, et al.
Published: (2022)
Why Do We Need Warm-up? A Theoretical Perspective
by: Alimisis, Foivos, et al.
Published: (2025)
by: Alimisis, Foivos, et al.
Published: (2025)
SDEs for Minimax Optimization
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
Beyond Universal Approximation Theorems: Algorithmic Uniform Approximation by Neural Networks Trained with Noisy Data
by: Kratsios, Anastasis, et al.
Published: (2025)
by: Kratsios, Anastasis, et al.
Published: (2025)
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023)
by: Francazi, Emanuele, et al.
Published: (2023)
Unbiased and Sign Compression in Distributed Learning: Comparing Noise Resilience via SDEs
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
On the Robustness of Neural Collapse and the Neural Collapse of Robustness
by: Su, Jingtong, et al.
Published: (2023)
by: Su, Jingtong, et al.
Published: (2023)
Gradient Scalability and Taylor Surrogation of Quantum Cost Landscapes
by: Meyer, Sabri, et al.
Published: (2025)
by: Meyer, Sabri, et al.
Published: (2025)
Double Momentum and Error Feedback for Clipping with Fast Rates and Differential Privacy
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
Where You Place the Norm Matters: From Prejudiced to Neutral Initializations
by: Francazi, Emanuele, et al.
Published: (2025)
by: Francazi, Emanuele, et al.
Published: (2025)
Adaptive Methods Are Preferable in High Privacy Settings: An SDE Perspective
by: Compagnoni, Enea Monzio, et al.
Published: (2026)
by: Compagnoni, Enea Monzio, et al.
Published: (2026)
When Bias Meets Trainability: Connecting Theories of Initialization
by: Bassi, Alberto, et al.
Published: (2025)
by: Bassi, Alberto, et al.
Published: (2025)
The Prevalence of Neural Collapse in Neural Multivariate Regression
by: Andriopoulos, George, et al.
Published: (2024)
by: Andriopoulos, George, et al.
Published: (2024)
A New Perspective To Understanding Multi-resolution Hash Encoding For Neural Fields
by: Luo, Steven Tin Sui
Published: (2025)
by: Luo, Steven Tin Sui
Published: (2025)
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
by: Anagnostidis, Sotiris, et al.
Published: (2023)
by: Anagnostidis, Sotiris, et al.
Published: (2023)
Missing Data Imputation using Neural Cellular Automata
by: Luu, Tin, et al.
Published: (2025)
by: Luu, Tin, et al.
Published: (2025)
Regret-Optimal Federated Transfer Learning for Kernel Regression with Applications in American Option Pricing
by: Yang, Xuwei, et al.
Published: (2023)
by: Yang, Xuwei, et al.
Published: (2023)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation
by: He, Xixiang, et al.
Published: (2026)
by: He, Xixiang, et al.
Published: (2026)
Efficient Matroid Bandit Linear Optimization Leveraging Unimodality
by: Delage, Aurélien, et al.
Published: (2025)
by: Delage, Aurélien, et al.
Published: (2025)
Cross Entropy versus Label Smoothing: A Neural Collapse Perspective
by: Guo, Li, et al.
Published: (2024)
by: Guo, Li, et al.
Published: (2024)
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
by: Súkeník, Peter, et al.
Published: (2024)
by: Súkeník, Peter, et al.
Published: (2024)
GReaTER: Generate Realistic Tabular data after data Enhancement and Reduction
by: Kwok, Tung Sum Thomas, et al.
Published: (2025)
by: Kwok, Tung Sum Thomas, et al.
Published: (2025)
EPA: Neural Collapse Inspired Robust Out-of-Distribution Detector
by: Zhang, Jiawei, et al.
Published: (2024)
by: Zhang, Jiawei, et al.
Published: (2024)
Rethinking Continual Learning with Progressive Neural Collapse
by: Wang, Zheng, et al.
Published: (2025)
by: Wang, Zheng, et al.
Published: (2025)
Pushing Boundaries: Mixup's Influence on Neural Collapse
by: Fisher, Quinn, et al.
Published: (2024)
by: Fisher, Quinn, et al.
Published: (2024)
Trojan Cleansing with Neural Collapse
by: Gu, Xihe, et al.
Published: (2024)
by: Gu, Xihe, et al.
Published: (2024)
Similar Items
-
Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization
by: Masarczyk, Wojciech, et al.
Published: (2025) -
Characterizing Overfitting in Kernel Ridgeless Regression Through the Eigenspectrum
by: Cheng, Tin Sum, et al.
Published: (2024) -
A Comprehensive Analysis on the Learning Curve in Kernel Ridge Regression
by: Cheng, Tin Sum, et al.
Published: (2024) -
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
by: Zhao, Jim, et al.
Published: (2024) -
Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters
by: Kratsios, Anastasis, et al.
Published: (2025)