Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Feng, Kunin, Daniel, Yamamura, Atsushi, Ganguli, Surya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
di: Kunin, Daniel, et al.
Pubblicazione: (2024)
di: Kunin, Daniel, et al.
Pubblicazione: (2024)
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
di: Yamamura, Atsushi, et al.
Pubblicazione: (2024)
di: Yamamura, Atsushi, et al.
Pubblicazione: (2024)
Stochastic Resetting Mitigates Latent Gradient Bias of SGD from Label Noise
di: Bae, Youngkyoung, et al.
Pubblicazione: (2024)
di: Bae, Youngkyoung, et al.
Pubblicazione: (2024)
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
di: Chen, Feng, et al.
Pubblicazione: (2025)
di: Chen, Feng, et al.
Pubblicazione: (2025)
DC-SGD: Differentially Private SGD with Dynamic Clipping through Gradient Norm Distribution Estimation
di: Wei, Chengkun, et al.
Pubblicazione: (2025)
di: Wei, Chengkun, et al.
Pubblicazione: (2025)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
di: Zhang, Tongcheng, et al.
Pubblicazione: (2026)
di: Zhang, Tongcheng, et al.
Pubblicazione: (2026)
Deriving Neural Scaling Laws from the statistics of natural language
di: Cagnetta, Francesco, et al.
Pubblicazione: (2026)
di: Cagnetta, Francesco, et al.
Pubblicazione: (2026)
Memory Constrained Dynamic Subnetwork Update for Transfer Learning
di: Quélennec, Aël, et al.
Pubblicazione: (2025)
di: Quélennec, Aël, et al.
Pubblicazione: (2025)
Continual Deep Learning on the Edge via Stochastic Local Competition among Subnetworks
di: Christophides, Theodoros, et al.
Pubblicazione: (2024)
di: Christophides, Theodoros, et al.
Pubblicazione: (2024)
Noise Balance and Stationary Distribution of Stochastic Gradient Descent
di: Ziyin, Liu, et al.
Pubblicazione: (2023)
di: Ziyin, Liu, et al.
Pubblicazione: (2023)
Instilling Inductive Biases with Subnetworks
di: Zhang, Enyan, et al.
Pubblicazione: (2023)
di: Zhang, Enyan, et al.
Pubblicazione: (2023)
Unveiling m-Sharpness Through the Structure of Stochastic Gradient Noise
di: Luo, Haocheng, et al.
Pubblicazione: (2025)
di: Luo, Haocheng, et al.
Pubblicazione: (2025)
From Kepler to Newton: Inductive Biases Guide Learned World Models in Transformers
di: Liu, Ziming, et al.
Pubblicazione: (2026)
di: Liu, Ziming, et al.
Pubblicazione: (2026)
Enhancing Output Diversity Improves Conjugate Gradient-based Adversarial Attacks
di: Yamamura, Keiichiro, et al.
Pubblicazione: (2024)
di: Yamamura, Keiichiro, et al.
Pubblicazione: (2024)
Model Parallelism With Subnetwork Data Parallelism
di: Singh, Vaibhav, et al.
Pubblicazione: (2025)
di: Singh, Vaibhav, et al.
Pubblicazione: (2025)
TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale
di: Ganguli, Anurup
Pubblicazione: (2026)
di: Ganguli, Anurup
Pubblicazione: (2026)
An analytic theory of creativity in convolutional diffusion models
di: Kamb, Mason, et al.
Pubblicazione: (2024)
di: Kamb, Mason, et al.
Pubblicazione: (2024)
FedSI: Federated Subnetwork Inference for Efficient Uncertainty Quantification
di: Chen, Hui, et al.
Pubblicazione: (2024)
di: Chen, Hui, et al.
Pubblicazione: (2024)
Contrastive Concept-Tree Search for LLM-Assisted Algorithm Discovery
di: Leleu, Timothee, et al.
Pubblicazione: (2026)
di: Leleu, Timothee, et al.
Pubblicazione: (2026)
Constrained Exploration via Reflected Replica Exchange Stochastic Gradient Langevin Dynamics
di: Zheng, Haoyang, et al.
Pubblicazione: (2024)
di: Zheng, Haoyang, et al.
Pubblicazione: (2024)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
di: Liao, Fangshuo, et al.
Pubblicazione: (2026)
di: Liao, Fangshuo, et al.
Pubblicazione: (2026)
The ODE Method for Stochastic Approximation and Reinforcement Learning with Markovian Noise
di: Liu, Shuze Daniel, et al.
Pubblicazione: (2024)
di: Liu, Shuze Daniel, et al.
Pubblicazione: (2024)
Enhancing DP-SGD through Non-monotonous Adaptive Scaling Gradient Weight
di: Huang, Tao, et al.
Pubblicazione: (2024)
di: Huang, Tao, et al.
Pubblicazione: (2024)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
di: Deora, Puneesh, et al.
Pubblicazione: (2025)
di: Deora, Puneesh, et al.
Pubblicazione: (2025)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
di: Yu, Dingzhi, et al.
Pubblicazione: (2026)
di: Yu, Dingzhi, et al.
Pubblicazione: (2026)
RQP-SGD: Differential Private Machine Learning through Noisy SGD and Randomized Quantization
di: Feng, Ce, et al.
Pubblicazione: (2024)
di: Feng, Ce, et al.
Pubblicazione: (2024)
Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective
di: Zhang, Jinouwen, et al.
Pubblicazione: (2024)
di: Zhang, Jinouwen, et al.
Pubblicazione: (2024)
A Combinatorial Theory of Dropout: Subnetworks, Graph Geometry, and Generalization
di: Dhayalkar, Sahil Rajesh
Pubblicazione: (2025)
di: Dhayalkar, Sahil Rajesh
Pubblicazione: (2025)
The Multiple Ticket Hypothesis: Random Sparse Subnetworks Suffice for RLVR
di: Adewuyi, Israel, et al.
Pubblicazione: (2026)
di: Adewuyi, Israel, et al.
Pubblicazione: (2026)
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
di: Thudi, Anvith, et al.
Pubblicazione: (2023)
di: Thudi, Anvith, et al.
Pubblicazione: (2023)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
di: Naganuma, Hiroki, et al.
Pubblicazione: (2026)
di: Naganuma, Hiroki, et al.
Pubblicazione: (2026)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
di: Bayazit, Deniz, et al.
Pubblicazione: (2023)
di: Bayazit, Deniz, et al.
Pubblicazione: (2023)
The Homogeneity Trap: Spectral Collapse in Doubly-Stochastic Deep Networks
di: Liu, Yizhi
Pubblicazione: (2026)
di: Liu, Yizhi
Pubblicazione: (2026)
Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models
di: Balashov, Andrii
Pubblicazione: (2025)
di: Balashov, Andrii
Pubblicazione: (2025)
TACTiS-2: Better, Faster, Simpler Attentional Copulas for Multivariate Time Series
di: Ashok, Arjun, et al.
Pubblicazione: (2023)
di: Ashok, Arjun, et al.
Pubblicazione: (2023)
Exploring Subnetwork Interactions in Heterogeneous Brain Network via Prior-Informed Graph Learning
di: Liu, Siyu, et al.
Pubblicazione: (2026)
di: Liu, Siyu, et al.
Pubblicazione: (2026)
Task-specific Subnetwork Discovery in Reinforcement Learning for Autonomous Underwater Navigation
di: Liu, Yi-Ling, et al.
Pubblicazione: (2026)
di: Liu, Yi-Ling, et al.
Pubblicazione: (2026)
Emergence of Globally Attracting Fixed Points in Deep Neural Networks With Nonlinear Activations
di: Joudaki, Amir, et al.
Pubblicazione: (2024)
di: Joudaki, Amir, et al.
Pubblicazione: (2024)
Almost Bayesian: The Fractal Dynamics of Stochastic Gradient Descent
di: Hennick, Max, et al.
Pubblicazione: (2025)
di: Hennick, Max, et al.
Pubblicazione: (2025)
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
di: Agarwal, Naman, et al.
Pubblicazione: (2025)
di: Agarwal, Naman, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
di: Kunin, Daniel, et al.
Pubblicazione: (2024) -
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
di: Yamamura, Atsushi, et al.
Pubblicazione: (2024) -
Stochastic Resetting Mitigates Latent Gradient Bias of SGD from Label Noise
di: Bae, Youngkyoung, et al.
Pubblicazione: (2024) -
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
di: Chen, Feng, et al.
Pubblicazione: (2025) -
DC-SGD: Differentially Private SGD with Dynamic Clipping through Gradient Norm Distribution Estimation
di: Wei, Chengkun, et al.
Pubblicazione: (2025)