From SGD to Spectra: A Theory of Neural Network Weight Dynamics
Fuente:
arXiv
Guardado en:
| Autores principales: | Olsen, Brian Richard, Fatehmanesh, Sam, Xiao, Frank, Kumarappan, Adarsh, Gajula, Anirudh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
por: Kumarappan, Adarsh, et al.
Publicado: (2025)
por: Kumarappan, Adarsh, et al.
Publicado: (2025)
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
por: Kumarappan, Adarsh, et al.
Publicado: (2025)
por: Kumarappan, Adarsh, et al.
Publicado: (2025)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
por: Galanti, Tomer, et al.
Publicado: (2022)
por: Galanti, Tomer, et al.
Publicado: (2022)
Sentiment-Aware Recommendation Systems in E-Commerce: A Review from a Natural Language Processing Perspective
por: Gajula, Yogesh
Publicado: (2025)
por: Gajula, Yogesh
Publicado: (2025)
LeanAgent: Lifelong Learning for Formal Theorem Proving
por: Kumarappan, Adarsh, et al.
Publicado: (2024)
por: Kumarappan, Adarsh, et al.
Publicado: (2024)
Optimal Condition for Initialization Variance in Deep Neural Networks: An SGD Dynamics Perspective
por: Horii, Hiroshi, et al.
Publicado: (2025)
por: Horii, Hiroshi, et al.
Publicado: (2025)
SGD with Partial Hessian for Deep Neural Networks Optimization
por: Sun, Ying, et al.
Publicado: (2024)
por: Sun, Ying, et al.
Publicado: (2024)
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
por: Sadrtdinov, Ildus, et al.
Publicado: (2025)
por: Sadrtdinov, Ildus, et al.
Publicado: (2025)
Memorization in Graph Neural Networks
por: Jamadandi, Adarsh, et al.
Publicado: (2025)
por: Jamadandi, Adarsh, et al.
Publicado: (2025)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
por: Marshall, Noah, et al.
Publicado: (2024)
por: Marshall, Noah, et al.
Publicado: (2024)
DP-SGD Without Clipping: The Lipschitz Neural Network Way
por: Bethune, Louis, et al.
Publicado: (2023)
por: Bethune, Louis, et al.
Publicado: (2023)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
por: Meterez, Alexandru, et al.
Publicado: (2025)
por: Meterez, Alexandru, et al.
Publicado: (2025)
Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed SGD
por: Wan, Yijun, et al.
Publicado: (2023)
por: Wan, Yijun, et al.
Publicado: (2023)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
por: Tanguy, Eloi
Publicado: (2023)
por: Tanguy, Eloi
Publicado: (2023)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
por: Beneventano, Pierfrancesco, et al.
Publicado: (2024)
por: Beneventano, Pierfrancesco, et al.
Publicado: (2024)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and Guidelines
por: Morad, Itai, et al.
Publicado: (2026)
por: Morad, Itai, et al.
Publicado: (2026)
Weight Spectra Induced Efficient Model Adaptation
por: Si, Chongjie, et al.
Publicado: (2025)
por: Si, Chongjie, et al.
Publicado: (2025)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
por: Zhang, Tongcheng, et al.
Publicado: (2026)
por: Zhang, Tongcheng, et al.
Publicado: (2026)
Diffusion-Based Neural Network Weights Generation
por: Soro, Bedionita, et al.
Publicado: (2024)
por: Soro, Bedionita, et al.
Publicado: (2024)
Numerical simulation of transient heat conduction with moving heat source using Physics Informed Neural Networks
por: Kalyan, Anirudh, et al.
Publicado: (2025)
por: Kalyan, Anirudh, et al.
Publicado: (2025)
From Weight Perturbation to Feature Attribution for Explaining Fully Connected Neural Networks
por: Lymperopoulos, Thodoris, et al.
Publicado: (2026)
por: Lymperopoulos, Thodoris, et al.
Publicado: (2026)
Weight Initialization and Variance Dynamics in Deep Neural Networks and Large Language Models
por: Han, Yankun
Publicado: (2025)
por: Han, Yankun
Publicado: (2025)
A Generalized Singular Value Theory for Neural Networks
por: Brown, Brian Charles, et al.
Publicado: (2026)
por: Brown, Brian Charles, et al.
Publicado: (2026)
Less is More: Efficient Weight Farcasting with 1-Layer Neural Network
por: Shou, Xiao, et al.
Publicado: (2025)
por: Shou, Xiao, et al.
Publicado: (2025)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
por: Xie, Shengping, et al.
Publicado: (2025)
por: Xie, Shengping, et al.
Publicado: (2025)
Balancing Utility and Privacy: Dynamically Private SGD with Random Projection
por: Jiang, Zhanhong, et al.
Publicado: (2025)
por: Jiang, Zhanhong, et al.
Publicado: (2025)
On the Stability of Nonlinear Dynamics in GD and SGD: Beyond Quadratic Potentials
por: Mulayoff, Rotem, et al.
Publicado: (2026)
por: Mulayoff, Rotem, et al.
Publicado: (2026)
Enhancing DP-SGD through Non-monotonous Adaptive Scaling Gradient Weight
por: Huang, Tao, et al.
Publicado: (2024)
por: Huang, Tao, et al.
Publicado: (2024)
Comparing Spectral Bias and Robustness For Two-Layer Neural Networks: SGD vs Adaptive Random Fourier Features
por: Kammonen, Aku, et al.
Publicado: (2024)
por: Kammonen, Aku, et al.
Publicado: (2024)
Global Convergence of SGD On Two Layer Neural Nets
por: Gopalani, Pulkit, et al.
Publicado: (2022)
por: Gopalani, Pulkit, et al.
Publicado: (2022)
Cooperative SGD with Dynamic Mixing Matrices
por: Sarkar, Soumya, et al.
Publicado: (2025)
por: Sarkar, Soumya, et al.
Publicado: (2025)
Accurate and Scalable Estimation of Epistemic Uncertainty for Graph Neural Networks
por: Trivedi, Puja, et al.
Publicado: (2024)
por: Trivedi, Puja, et al.
Publicado: (2024)
Deep Neural Network for Phonon-Assisted Optical Spectra in Semiconductors
por: Gu, Qiangqiang, et al.
Publicado: (2025)
por: Gu, Qiangqiang, et al.
Publicado: (2025)
Spectra-Guided Neural Tucker Factorization
por: Wang, Fusheng, et al.
Publicado: (2026)
por: Wang, Fusheng, et al.
Publicado: (2026)
From Privacy to Generalization: Linear Max-Information Bounds for DP-SGD
por: Lampert, Christoph H., et al.
Publicado: (2026)
por: Lampert, Christoph H., et al.
Publicado: (2026)
Style-based Clustering of Visual Artworks and the Play of Neural Style-Representations
por: Dangeti, Abhishek, et al.
Publicado: (2024)
por: Dangeti, Abhishek, et al.
Publicado: (2024)
From Gradient Clipping to Normalization for Heavy Tailed SGD
por: Hübler, Florian, et al.
Publicado: (2024)
por: Hübler, Florian, et al.
Publicado: (2024)
Signal Processing Meets SGD: From Momentum to Filter
por: Yao, Zhipeng, et al.
Publicado: (2023)
por: Yao, Zhipeng, et al.
Publicado: (2023)
Ejemplares similares
-
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
por: Kumarappan, Adarsh, et al.
Publicado: (2025) -
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
por: Kumarappan, Adarsh, et al.
Publicado: (2025) -
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
por: Kumarappan, Adarsh, et al.
Publicado: (2026) -
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
por: Galanti, Tomer, et al.
Publicado: (2022) -
Sentiment-Aware Recommendation Systems in E-Commerce: A Review from a Natural Language Processing Perspective
por: Gajula, Yogesh
Publicado: (2025)