Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
Fuente:
arXiv
Saved in:
| Main Authors: | Beaglehole, Daniel, Mitliagkas, Ioannis, Agarwala, Atish |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Interplay Between Stepsize Tuning and Progressive Sharpening
by: Roulet, Vincent, et al.
Published: (2023)
by: Roulet, Vincent, et al.
Published: (2023)
Towards efficient representation identification in supervised learning
by: Ahuja, Kartik, et al.
Published: (2022)
by: Ahuja, Kartik, et al.
Published: (2022)
Per-example gradients: a new frontier for understanding and improving optimizers
by: Roulet, Vincent, et al.
Published: (2025)
by: Roulet, Vincent, et al.
Published: (2025)
The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks
by: Cha, Taehun, et al.
Published: (2026)
by: Cha, Taehun, et al.
Published: (2026)
Fractional-order spike-timing-dependent gradient descent for multi-layer spiking neural networks
by: Yang, Yi, et al.
Published: (2024)
by: Yang, Yi, et al.
Published: (2024)
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
by: Naganuma, Hiroki, et al.
Published: (2023)
by: Naganuma, Hiroki, et al.
Published: (2023)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)
by: Han, Qiyang, et al.
Published: (2025)
Are aligned neural networks adversarially aligned?
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation
by: Mahajan, Divyat, et al.
Published: (2022)
by: Mahajan, Divyat, et al.
Published: (2022)
High dimensional theory of two-phase optimizers
by: Agarwala, Atish
Published: (2026)
by: Agarwala, Atish
Published: (2026)
Performance Control in Early Exiting to Deploy Large Models at the Same Cost of Smaller Ones
by: Mofakhami, Mehrnaz, et al.
Published: (2024)
by: Mofakhami, Mehrnaz, et al.
Published: (2024)
State-space models can learn in-context by gradient descent
by: Sushma, Neeraj Mohan, et al.
Published: (2024)
by: Sushma, Neeraj Mohan, et al.
Published: (2024)
Compositional Risk Minimization
by: Mahajan, Divyat, et al.
Published: (2024)
by: Mahajan, Divyat, et al.
Published: (2024)
ProtAlign: Contrastive learning paradigm for Sequence and structure alignment
by: Ranganath, Aditya, et al.
Published: (2026)
by: Ranganath, Aditya, et al.
Published: (2026)
Evaluating alignment between humans and neural network representations in image-based learning tasks
by: Demircan, Can, et al.
Published: (2023)
by: Demircan, Can, et al.
Published: (2023)
Navigating Potholes with Geometry-Aware Sharpness Minimization
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
Graph neural networks and non-commuting operators
by: Velasco, Mauricio, et al.
Published: (2024)
by: Velasco, Mauricio, et al.
Published: (2024)
Efficient line search for optimizing Area Under the ROC Curve in gradient descent
by: Fowler, Jadon, et al.
Published: (2024)
by: Fowler, Jadon, et al.
Published: (2024)
Learning to Defer for Causal Discovery with Imperfect Experts
by: Clivio, Oscar, et al.
Published: (2025)
by: Clivio, Oscar, et al.
Published: (2025)
$σ$-PCA: a building block for neural learning of identifiable linear transformations
by: Kanavati, Fahdi, et al.
Published: (2023)
by: Kanavati, Fahdi, et al.
Published: (2023)
Steering Autoregressive Music Generation with Recursive Feature Machines
by: Zhao, Daniel, et al.
Published: (2025)
by: Zhao, Daniel, et al.
Published: (2025)
Human alignment of neural network representations
by: Muttenthaler, Lukas, et al.
Published: (2022)
by: Muttenthaler, Lukas, et al.
Published: (2022)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
by: Zucchet, Nicolas, et al.
Published: (2024)
by: Zucchet, Nicolas, et al.
Published: (2024)
Adaptive multiple optimal learning factors for neural network training
by: Challagundla, Jeshwanth
Published: (2024)
by: Challagundla, Jeshwanth
Published: (2024)
Understanding the learned look-ahead behavior of chess neural networks
by: Cruz, Diogo
Published: (2025)
by: Cruz, Diogo
Published: (2025)
Do graph neural network states contain graph properties?
by: Pelletreau-Duris, Tom, et al.
Published: (2024)
by: Pelletreau-Duris, Tom, et al.
Published: (2024)
Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
by: Mahajan, Divyat, et al.
Published: (2025)
by: Mahajan, Divyat, et al.
Published: (2025)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
When majority rules, minority loses: bias amplification of gradient descent
by: Bachoc, François, et al.
Published: (2025)
by: Bachoc, François, et al.
Published: (2025)
Toward universal steering and monitoring of AI models
by: Beaglehole, Daniel, et al.
Published: (2025)
by: Beaglehole, Daniel, et al.
Published: (2025)
Average gradient outer product as a mechanism for deep neural collapse
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
by: Boix-Adsera, Enric, et al.
Published: (2025)
by: Boix-Adsera, Enric, et al.
Published: (2025)
The duality structure gradient descent algorithm: analysis and applications to neural networks
by: Flynn, Thomas
Published: (2017)
by: Flynn, Thomas
Published: (2017)
Computing the gradients with respect to all parameters of a quantum neural network using a single circuit
by: He, Guang Ping
Published: (2023)
by: He, Guang Ping
Published: (2023)
Feature contamination: Neural networks learn uncorrelated features and fail to generalize
by: Zhang, Tianren, et al.
Published: (2024)
by: Zhang, Tianren, et al.
Published: (2024)
Binary structured physics-informed neural networks for solving equations with rapidly changing solutions
by: Liu, Yanzhi, et al.
Published: (2024)
by: Liu, Yanzhi, et al.
Published: (2024)
The ecosystem of machine learning competitions: Platforms, participants, and their impact on AI development
by: Nasios, Ioannis
Published: (2026)
by: Nasios, Ioannis
Published: (2026)
Investigating potential causes of Sepsis with Bayesian network structure learning
by: Petrungaro, Bruno, et al.
Published: (2024)
by: Petrungaro, Bruno, et al.
Published: (2024)
Dimensions underlying the representational alignment of deep neural networks with humans
by: Mahner, Florian P., et al.
Published: (2024)
by: Mahner, Florian P., et al.
Published: (2024)
Representations learnt by SGD and Adaptive learning rules: Conditions that vary sparsity and selectivity in neural networks
by: Park, Jin Hyun
Published: (2022)
by: Park, Jin Hyun
Published: (2022)
Similar Items
-
On the Interplay Between Stepsize Tuning and Progressive Sharpening
by: Roulet, Vincent, et al.
Published: (2023) -
Towards efficient representation identification in supervised learning
by: Ahuja, Kartik, et al.
Published: (2022) -
Per-example gradients: a new frontier for understanding and improving optimizers
by: Roulet, Vincent, et al.
Published: (2025) -
The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks
by: Cha, Taehun, et al.
Published: (2026) -
Fractional-order spike-timing-dependent gradient descent for multi-layer spiking neural networks
by: Yang, Yi, et al.
Published: (2024)