Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed SGD
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Yijun, Barsbey, Melih, Zaidi, Abdellatif, Simsekli, Umut |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Algorithmic Stability of Stochastic Gradient Descent with Momentum under Heavy-Tailed Noise
by: Dang, Thanh, et al.
Published: (2025)
by: Dang, Thanh, et al.
Published: (2025)
On the Interaction of Compressibility and Adversarial Robustness
by: Barsbey, Melih, et al.
Published: (2025)
by: Barsbey, Melih, et al.
Published: (2025)
Heavy-Tail Phenomenon in Decentralized SGD
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
Privacy of SGD under Gaussian or Heavy-Tailed Noise: Guarantees without Gradient Clipping
by: Şimşekli, Umut, et al.
Published: (2024)
by: Şimşekli, Umut, et al.
Published: (2024)
Generalization Bounds for Heavy-Tailed SDEs through the Fractional Fokker-Planck Equation
by: Dupuis, Benjamin, et al.
Published: (2024)
by: Dupuis, Benjamin, et al.
Published: (2024)
Symmetries in Overparametrized Neural Networks: A Mean-Field View
by: Maass, Javier, et al.
Published: (2024)
by: Maass, Javier, et al.
Published: (2024)
Heavy-Tailed Diffusion with Denoising Lévy Probabilistic Models
by: Shariatian, Dario, et al.
Published: (2024)
by: Shariatian, Dario, et al.
Published: (2024)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
by: Tanguy, Eloi
Published: (2023)
by: Tanguy, Eloi
Published: (2023)
Rényi Differential Privacy for Heavy-Tailed SDEs via Fractional Poincaré Inequalities
by: Dupuis, Benjamin, et al.
Published: (2025)
by: Dupuis, Benjamin, et al.
Published: (2025)
Global Dynamics of Heavy-Tailed SGDs in Nonconvex Loss Landscape: Characterization and Control
by: Wang, Xingyu, et al.
Published: (2025)
by: Wang, Xingyu, et al.
Published: (2025)
Diffusion Models with Heavy-Tailed Targets: Score Estimation and Sampling Guarantees
by: Yu, Yifeng, et al.
Published: (2026)
by: Yu, Yifeng, et al.
Published: (2026)
Concentration of General Stochastic Approximation Under Heavy-Tailed Markovian Noise
by: Agrawal, Shubhada, et al.
Published: (2026)
by: Agrawal, Shubhada, et al.
Published: (2026)
Large Deviation Upper Bounds and Improved MSE Rates of Nonlinear SGD: Heavy-tailed Noise and Power of Symmetry
by: Armacki, Aleksandar, et al.
Published: (2024)
by: Armacki, Aleksandar, et al.
Published: (2024)
Lessons from Generalization Error Analysis of Federated Learning: You May Communicate Less Often!
by: Sefidgaran, Milad, et al.
Published: (2023)
by: Sefidgaran, Milad, et al.
Published: (2023)
Error estimates between SGD with momentum and underdamped Langevin diffusion
by: Guillin, Arnaud, et al.
Published: (2024)
by: Guillin, Arnaud, et al.
Published: (2024)
Data-dependent Generalization Bounds via Variable-Size Compressibility
by: Sefidgaran, Milad, et al.
Published: (2023)
by: Sefidgaran, Milad, et al.
Published: (2023)
Convergence, Sticking and Escape: Stochastic Dynamics Near Critical Points in SGD
by: Dudukalov, Dmitry, et al.
Published: (2025)
by: Dudukalov, Dmitry, et al.
Published: (2025)
Effective continuous equations for adaptive SGD: a stochastic analysis view
by: Callisti, Luca, et al.
Published: (2025)
by: Callisti, Luca, et al.
Published: (2025)
Training Overparametrized Neural Networks in Sublinear Time
by: Deng, Yichuan, et al.
Published: (2022)
by: Deng, Yichuan, et al.
Published: (2022)
Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and Compressibility
by: Barsbey, Melih, et al.
Published: (2025)
by: Barsbey, Melih, et al.
Published: (2025)
Diffusion Processes on Implicit Manifolds
by: Kawasaki-Borruat, Victor, et al.
Published: (2026)
by: Kawasaki-Borruat, Victor, et al.
Published: (2026)
Can SGD Handle Heavy-Tailed Noise?
by: Fatkhullin, Ilyas, et al.
Published: (2025)
by: Fatkhullin, Ilyas, et al.
Published: (2025)
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
by: Zhang, Huiming, et al.
Published: (2026)
by: Zhang, Huiming, et al.
Published: (2026)
Dropout Neural Network Training Viewed from a Percolation Perspective
by: Devlin, Finley, et al.
Published: (2025)
by: Devlin, Finley, et al.
Published: (2025)
Random Normed k-Means: A Paradigm-Shift in Clustering within Probabilistic Metric Spaces
by: Hemdanou, Abderrafik Laakel, et al.
Published: (2025)
by: Hemdanou, Abderrafik Laakel, et al.
Published: (2025)
Sharper Perturbed-Kullback-Leibler Exponential Tail Bounds for Beta and Dirichlet Distributions
by: Perrault, Pierre
Published: (2025)
by: Perrault, Pierre
Published: (2025)
Steady-State Behavior of Constant-Stepsize Stochastic Approximation: Gaussian Approximation and Tail Bounds
by: Wang, Zedong, et al.
Published: (2026)
by: Wang, Zedong, et al.
Published: (2026)
Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
From Gradient Clipping to Normalization for Heavy Tailed SGD
by: Hübler, Florian, et al.
Published: (2024)
by: Hübler, Florian, et al.
Published: (2024)
Partially Stochastic Infinitely Deep Bayesian Neural Networks
by: Calvo-Ordonez, Sergio, et al.
Published: (2024)
by: Calvo-Ordonez, Sergio, et al.
Published: (2024)
On the Epistemic Uncertainty of Overparametrized Neural Networks
by: Rügamer, David
Published: (2026)
by: Rügamer, David
Published: (2026)
Efficient Training of Deep Neural Operator Networks via Randomized Sampling
by: Karumuri, Sharmila, et al.
Published: (2024)
by: Karumuri, Sharmila, et al.
Published: (2024)
Explicit Density Approximation for Neural Implicit Samplers Using a Bernstein-Based Convex Divergence
by: de Frutos, José Manuel, et al.
Published: (2025)
by: de Frutos, José Manuel, et al.
Published: (2025)
Depth Degeneracy in Neural Networks: Vanishing Angles in Fully Connected ReLU Networks on Initialization
by: Jakub, Cameron, et al.
Published: (2023)
by: Jakub, Cameron, et al.
Published: (2023)
Convergence Analysis of Newton's Method for Neural Networks in the Overparameterized Limit
by: Riedl, Konstantin, et al.
Published: (2026)
by: Riedl, Konstantin, et al.
Published: (2026)
Random ReLU Neural Networks as Non-Gaussian Processes
by: Parhi, Rahul, et al.
Published: (2024)
by: Parhi, Rahul, et al.
Published: (2024)
Stochastic Port-Hamiltonian Neural Networks: Universal Approximation with Passivity Guarantees
by: Di Persio, Luca, et al.
Published: (2026)
by: Di Persio, Luca, et al.
Published: (2026)
Universality in Deep Neural Networks: An approach via the Lindeberg exchange principle
by: Giovagnini, Filippo, et al.
Published: (2026)
by: Giovagnini, Filippo, et al.
Published: (2026)
Exact Gradients for Stochastic Spiking Neural Networks Driven by Rough Signals
by: Holberg, Christian, et al.
Published: (2024)
by: Holberg, Christian, et al.
Published: (2024)
Deep Neural Networks as Iterated Function Systems and a Generalization Bound
by: Vacher, Jonathan
Published: (2026)
by: Vacher, Jonathan
Published: (2026)
Similar Items
-
Algorithmic Stability of Stochastic Gradient Descent with Momentum under Heavy-Tailed Noise
by: Dang, Thanh, et al.
Published: (2025) -
On the Interaction of Compressibility and Adversarial Robustness
by: Barsbey, Melih, et al.
Published: (2025) -
Heavy-Tail Phenomenon in Decentralized SGD
by: Gurbuzbalaban, Mert, et al.
Published: (2022) -
Privacy of SGD under Gaussian or Heavy-Tailed Noise: Guarantees without Gradient Clipping
by: Şimşekli, Umut, et al.
Published: (2024) -
Generalization Bounds for Heavy-Tailed SDEs through the Fractional Fokker-Planck Equation
by: Dupuis, Benjamin, et al.
Published: (2024)