Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Feng, Kunin, Daniel, Yamamura, Atsushi, Ganguli, Surya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
von: Kunin, Daniel, et al.
Veröffentlicht: (2024)
von: Kunin, Daniel, et al.
Veröffentlicht: (2024)
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
von: Yamamura, Atsushi, et al.
Veröffentlicht: (2024)
von: Yamamura, Atsushi, et al.
Veröffentlicht: (2024)
Stochastic Resetting Mitigates Latent Gradient Bias of SGD from Label Noise
von: Bae, Youngkyoung, et al.
Veröffentlicht: (2024)
von: Bae, Youngkyoung, et al.
Veröffentlicht: (2024)
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
von: Chen, Feng, et al.
Veröffentlicht: (2025)
von: Chen, Feng, et al.
Veröffentlicht: (2025)
DC-SGD: Differentially Private SGD with Dynamic Clipping through Gradient Norm Distribution Estimation
von: Wei, Chengkun, et al.
Veröffentlicht: (2025)
von: Wei, Chengkun, et al.
Veröffentlicht: (2025)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
von: Zhang, Tongcheng, et al.
Veröffentlicht: (2026)
von: Zhang, Tongcheng, et al.
Veröffentlicht: (2026)
Deriving Neural Scaling Laws from the statistics of natural language
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2026)
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2026)
Memory Constrained Dynamic Subnetwork Update for Transfer Learning
von: Quélennec, Aël, et al.
Veröffentlicht: (2025)
von: Quélennec, Aël, et al.
Veröffentlicht: (2025)
Continual Deep Learning on the Edge via Stochastic Local Competition among Subnetworks
von: Christophides, Theodoros, et al.
Veröffentlicht: (2024)
von: Christophides, Theodoros, et al.
Veröffentlicht: (2024)
Noise Balance and Stationary Distribution of Stochastic Gradient Descent
von: Ziyin, Liu, et al.
Veröffentlicht: (2023)
von: Ziyin, Liu, et al.
Veröffentlicht: (2023)
Instilling Inductive Biases with Subnetworks
von: Zhang, Enyan, et al.
Veröffentlicht: (2023)
von: Zhang, Enyan, et al.
Veröffentlicht: (2023)
Unveiling m-Sharpness Through the Structure of Stochastic Gradient Noise
von: Luo, Haocheng, et al.
Veröffentlicht: (2025)
von: Luo, Haocheng, et al.
Veröffentlicht: (2025)
From Kepler to Newton: Inductive Biases Guide Learned World Models in Transformers
von: Liu, Ziming, et al.
Veröffentlicht: (2026)
von: Liu, Ziming, et al.
Veröffentlicht: (2026)
Enhancing Output Diversity Improves Conjugate Gradient-based Adversarial Attacks
von: Yamamura, Keiichiro, et al.
Veröffentlicht: (2024)
von: Yamamura, Keiichiro, et al.
Veröffentlicht: (2024)
Model Parallelism With Subnetwork Data Parallelism
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale
von: Ganguli, Anurup
Veröffentlicht: (2026)
von: Ganguli, Anurup
Veröffentlicht: (2026)
An analytic theory of creativity in convolutional diffusion models
von: Kamb, Mason, et al.
Veröffentlicht: (2024)
von: Kamb, Mason, et al.
Veröffentlicht: (2024)
FedSI: Federated Subnetwork Inference for Efficient Uncertainty Quantification
von: Chen, Hui, et al.
Veröffentlicht: (2024)
von: Chen, Hui, et al.
Veröffentlicht: (2024)
Contrastive Concept-Tree Search for LLM-Assisted Algorithm Discovery
von: Leleu, Timothee, et al.
Veröffentlicht: (2026)
von: Leleu, Timothee, et al.
Veröffentlicht: (2026)
Constrained Exploration via Reflected Replica Exchange Stochastic Gradient Langevin Dynamics
von: Zheng, Haoyang, et al.
Veröffentlicht: (2024)
von: Zheng, Haoyang, et al.
Veröffentlicht: (2024)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
von: Liao, Fangshuo, et al.
Veröffentlicht: (2026)
von: Liao, Fangshuo, et al.
Veröffentlicht: (2026)
The ODE Method for Stochastic Approximation and Reinforcement Learning with Markovian Noise
von: Liu, Shuze Daniel, et al.
Veröffentlicht: (2024)
von: Liu, Shuze Daniel, et al.
Veröffentlicht: (2024)
Enhancing DP-SGD through Non-monotonous Adaptive Scaling Gradient Weight
von: Huang, Tao, et al.
Veröffentlicht: (2024)
von: Huang, Tao, et al.
Veröffentlicht: (2024)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
von: Deora, Puneesh, et al.
Veröffentlicht: (2025)
von: Deora, Puneesh, et al.
Veröffentlicht: (2025)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
RQP-SGD: Differential Private Machine Learning through Noisy SGD and Randomized Quantization
von: Feng, Ce, et al.
Veröffentlicht: (2024)
von: Feng, Ce, et al.
Veröffentlicht: (2024)
Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective
von: Zhang, Jinouwen, et al.
Veröffentlicht: (2024)
von: Zhang, Jinouwen, et al.
Veröffentlicht: (2024)
A Combinatorial Theory of Dropout: Subnetworks, Graph Geometry, and Generalization
von: Dhayalkar, Sahil Rajesh
Veröffentlicht: (2025)
von: Dhayalkar, Sahil Rajesh
Veröffentlicht: (2025)
The Multiple Ticket Hypothesis: Random Sparse Subnetworks Suffice for RLVR
von: Adewuyi, Israel, et al.
Veröffentlicht: (2026)
von: Adewuyi, Israel, et al.
Veröffentlicht: (2026)
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
von: Thudi, Anvith, et al.
Veröffentlicht: (2023)
von: Thudi, Anvith, et al.
Veröffentlicht: (2023)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
von: Naganuma, Hiroki, et al.
Veröffentlicht: (2026)
von: Naganuma, Hiroki, et al.
Veröffentlicht: (2026)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
von: Bayazit, Deniz, et al.
Veröffentlicht: (2023)
von: Bayazit, Deniz, et al.
Veröffentlicht: (2023)
The Homogeneity Trap: Spectral Collapse in Doubly-Stochastic Deep Networks
von: Liu, Yizhi
Veröffentlicht: (2026)
von: Liu, Yizhi
Veröffentlicht: (2026)
Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models
von: Balashov, Andrii
Veröffentlicht: (2025)
von: Balashov, Andrii
Veröffentlicht: (2025)
TACTiS-2: Better, Faster, Simpler Attentional Copulas for Multivariate Time Series
von: Ashok, Arjun, et al.
Veröffentlicht: (2023)
von: Ashok, Arjun, et al.
Veröffentlicht: (2023)
Exploring Subnetwork Interactions in Heterogeneous Brain Network via Prior-Informed Graph Learning
von: Liu, Siyu, et al.
Veröffentlicht: (2026)
von: Liu, Siyu, et al.
Veröffentlicht: (2026)
Task-specific Subnetwork Discovery in Reinforcement Learning for Autonomous Underwater Navigation
von: Liu, Yi-Ling, et al.
Veröffentlicht: (2026)
von: Liu, Yi-Ling, et al.
Veröffentlicht: (2026)
Emergence of Globally Attracting Fixed Points in Deep Neural Networks With Nonlinear Activations
von: Joudaki, Amir, et al.
Veröffentlicht: (2024)
von: Joudaki, Amir, et al.
Veröffentlicht: (2024)
Almost Bayesian: The Fractal Dynamics of Stochastic Gradient Descent
von: Hennick, Max, et al.
Veröffentlicht: (2025)
von: Hennick, Max, et al.
Veröffentlicht: (2025)
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
von: Agarwal, Naman, et al.
Veröffentlicht: (2025)
von: Agarwal, Naman, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
von: Kunin, Daniel, et al.
Veröffentlicht: (2024) -
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
von: Yamamura, Atsushi, et al.
Veröffentlicht: (2024) -
Stochastic Resetting Mitigates Latent Gradient Bias of SGD from Label Noise
von: Bae, Youngkyoung, et al.
Veröffentlicht: (2024) -
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
von: Chen, Feng, et al.
Veröffentlicht: (2025) -
DC-SGD: Differentially Private SGD with Dynamic Clipping through Gradient Norm Distribution Estimation
von: Wei, Chengkun, et al.
Veröffentlicht: (2025)