Average gradient outer product as a mechanism for deep neural collapse
Fuente:
arXiv
Guardado en:
| Autores principales: | Beaglehole, Daniel, Súkeník, Peter, Mondelli, Marco, Belkin, Mikhail |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
por: Mallinar, Neil, et al.
Publicado: (2024)
por: Mallinar, Neil, et al.
Publicado: (2024)
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
por: Súkeník, Peter, et al.
Publicado: (2024)
por: Súkeník, Peter, et al.
Publicado: (2024)
Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers
por: Súkeník, Peter, et al.
Publicado: (2025)
por: Súkeník, Peter, et al.
Publicado: (2025)
Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention
por: Súkeník, Peter, et al.
Publicado: (2026)
por: Súkeník, Peter, et al.
Publicado: (2026)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
por: Jacot, Arthur, et al.
Publicado: (2024)
por: Jacot, Arthur, et al.
Publicado: (2024)
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
por: Beaglehole, Daniel, et al.
Publicado: (2025)
por: Beaglehole, Daniel, et al.
Publicado: (2025)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
por: Beaglehole, Daniel, et al.
Publicado: (2024)
por: Beaglehole, Daniel, et al.
Publicado: (2024)
Toward universal steering and monitoring of AI models
por: Beaglehole, Daniel, et al.
Publicado: (2025)
por: Beaglehole, Daniel, et al.
Publicado: (2025)
Neural Collapse Beyond the Unconstrained Features Model: Landscape, Dynamics, and Generalization in the Mean-Field Regime
por: Wu, Diyuan, et al.
Publicado: (2025)
por: Wu, Diyuan, et al.
Publicado: (2025)
Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-Parameterization
por: Bombari, Simone, et al.
Publicado: (2025)
por: Bombari, Simone, et al.
Publicado: (2025)
How Spurious Features Are Memorized: Precise Analysis for Random and NTK Features
por: Bombari, Simone, et al.
Publicado: (2023)
por: Bombari, Simone, et al.
Publicado: (2023)
Quadratic models for understanding catapult dynamics of neural networks
por: Zhu, Libin, et al.
Publicado: (2022)
por: Zhu, Libin, et al.
Publicado: (2022)
Privacy for Free in the Overparameterized Regime
por: Bombari, Simone, et al.
Publicado: (2024)
por: Bombari, Simone, et al.
Publicado: (2024)
Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features
por: Bombari, Simone, et al.
Publicado: (2024)
por: Bombari, Simone, et al.
Publicado: (2024)
Intriguing Properties of Input-dependent Randomized Smoothing
por: Súkeník, Peter, et al.
Publicado: (2021)
por: Súkeník, Peter, et al.
Publicado: (2021)
Optimal Regularization for Performative Learning
por: Cyffers, Edwige, et al.
Publicado: (2025)
por: Cyffers, Edwige, et al.
Publicado: (2025)
Matrix Denoising with Doubly Heteroscedastic Noise: Fundamental Limits and Optimal Spectral Methods
por: Zhang, Yihan, et al.
Publicado: (2024)
por: Zhang, Yihan, et al.
Publicado: (2024)
Linear Recursive Feature Machines provably recover low-rank matrices
por: Radhakrishnan, Adityanarayanan, et al.
Publicado: (2024)
por: Radhakrishnan, Adityanarayanan, et al.
Publicado: (2024)
Improved Convergence of Score-Based Diffusion Models via Prediction-Correction
por: Pedrotti, Francesco, et al.
Publicado: (2023)
por: Pedrotti, Francesco, et al.
Publicado: (2023)
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
por: Mirtaheri, Parsa, et al.
Publicado: (2026)
por: Mirtaheri, Parsa, et al.
Publicado: (2026)
The role of class encoding in neural collapse
por: Massion, Bastien, et al.
Publicado: (2026)
por: Massion, Bastien, et al.
Publicado: (2026)
High-Dimensional Private Linear Regression with Optimal Rates
por: Bombari, Simone, et al.
Publicado: (2025)
por: Bombari, Simone, et al.
Publicado: (2025)
A Law of Data Reconstruction for Random Features (and Beyond)
por: Iurada, Leonardo, et al.
Publicado: (2025)
por: Iurada, Leonardo, et al.
Publicado: (2025)
Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression
por: Wu, Diyuan, et al.
Publicado: (2026)
por: Wu, Diyuan, et al.
Publicado: (2026)
High-dimensional Analysis of Synthetic Data Selection
por: Rezaei, Parham, et al.
Publicado: (2025)
por: Rezaei, Parham, et al.
Publicado: (2025)
Precise Asymptotics for Spectral Methods in Mixed Generalized Linear Models
por: Zhang, Yihan, et al.
Publicado: (2022)
por: Zhang, Yihan, et al.
Publicado: (2022)
Mirror Descent on Reproducing Kernel Banach Spaces
por: Kumar, Akash, et al.
Publicado: (2024)
por: Kumar, Akash, et al.
Publicado: (2024)
General and Efficient Steering of Unconditional Diffusion
por: Wang, Qingsong, et al.
Publicado: (2026)
por: Wang, Qingsong, et al.
Publicado: (2026)
A Gap Between the Gaussian RKHS and Neural Networks: An Infinite-Center Asymptotic Analysis
por: Kumar, Akash, et al.
Publicado: (2025)
por: Kumar, Akash, et al.
Publicado: (2025)
Convergence of continuous-time stochastic gradient descent with applications to deep neural networks
por: Lugosi, Gabor, et al.
Publicado: (2024)
por: Lugosi, Gabor, et al.
Publicado: (2024)
Spectral Estimators for Multi-Index Models: Precise Asymptotics and Optimal Weak Recovery
por: Kovačević, Filip, et al.
Publicado: (2025)
por: Kovačević, Filip, et al.
Publicado: (2025)
Convergence of stochastic gradient descent under a local Lojasiewicz condition for deep neural networks
por: An, Jing, et al.
Publicado: (2023)
por: An, Jing, et al.
Publicado: (2023)
Fast training of large kernel models with delayed projections
por: Abedsoltan, Amirhesam, et al.
Publicado: (2024)
por: Abedsoltan, Amirhesam, et al.
Publicado: (2024)
Context-Scaling versus Task-Scaling in In-Context Learning
por: Abedsoltan, Amirhesam, et al.
Publicado: (2024)
por: Abedsoltan, Amirhesam, et al.
Publicado: (2024)
On the Nystrom Approximation for Preconditioning in Kernel Machines
por: Abedsoltan, Amirhesam, et al.
Publicado: (2023)
por: Abedsoltan, Amirhesam, et al.
Publicado: (2023)
Breaking Data Symmetry is Needed For Generalization in Feature Learning Kernels
por: Bernal, Marcel Tomàs, et al.
Publicado: (2026)
por: Bernal, Marcel Tomàs, et al.
Publicado: (2026)
The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks
por: Cha, Taehun, et al.
Publicado: (2026)
por: Cha, Taehun, et al.
Publicado: (2026)
Compression of Structured Data with Autoencoders: Provable Benefit of Nonlinearities and Depth
por: Kögler, Kevin, et al.
Publicado: (2024)
por: Kögler, Kevin, et al.
Publicado: (2024)
Attention with Trained Embeddings Provably Selects Important Tokens
por: Wu, Diyuan, et al.
Publicado: (2025)
por: Wu, Diyuan, et al.
Publicado: (2025)
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
por: Njaradi, Valentina, et al.
Publicado: (2026)
por: Njaradi, Valentina, et al.
Publicado: (2026)
Ejemplares similares
-
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
por: Mallinar, Neil, et al.
Publicado: (2024) -
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
por: Súkeník, Peter, et al.
Publicado: (2024) -
Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers
por: Súkeník, Peter, et al.
Publicado: (2025) -
Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention
por: Súkeník, Peter, et al.
Publicado: (2026) -
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
por: Jacot, Arthur, et al.
Publicado: (2024)