Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Xuan, Zhang, Han, Cao, Yuan, Zou, Difan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
by: Tang, Xuan, et al.
Published: (2025)
by: Tang, Xuan, et al.
Published: (2025)
Scaling Laws for Precision in High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2026)
by: Zhang, Dechen, et al.
Published: (2026)
Learning under Quantization for High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2025)
by: Zhang, Dechen, et al.
Published: (2025)
The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient
by: Li, Jichu, et al.
Published: (2026)
by: Li, Jichu, et al.
Published: (2026)
The Implicit Bias of Adam on Separable Data
by: Zhang, Chenyang, et al.
Published: (2024)
by: Zhang, Chenyang, et al.
Published: (2024)
Understanding the Dynamics of Demonstration Conflict in In-Context Learning
by: Jiao, Difan, et al.
Published: (2026)
by: Jiao, Difan, et al.
Published: (2026)
Gradient Descent Robustly Learns the Intrinsic Dimension of Data in Training Convolutional Neural Networks
by: Zhang, Chenyang, et al.
Published: (2025)
by: Zhang, Chenyang, et al.
Published: (2025)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
by: Chen, Xingwu, et al.
Published: (2025)
by: Chen, Xingwu, et al.
Published: (2025)
Structured Role-Aware Policy Optimization for Multimodal Reasoning
by: Jiang, Bingqing, et al.
Published: (2026)
by: Jiang, Bingqing, et al.
Published: (2026)
On the Memorization of Consistency Distillation for Diffusion Models
by: Jiang, Bingqing, et al.
Published: (2026)
by: Jiang, Bingqing, et al.
Published: (2026)
On the Robustness of Transformers against Context Hijacking for Linear Classification
by: Li, Tianle, et al.
Published: (2025)
by: Li, Tianle, et al.
Published: (2025)
Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Kernel Regression in Structured Non-IID Settings: Theory and Implications for Denoising Score Learning
by: Zhang, Dechen, et al.
Published: (2025)
by: Zhang, Dechen, et al.
Published: (2025)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
by: Xie, Chengxing, et al.
Published: (2024)
by: Xie, Chengxing, et al.
Published: (2024)
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
by: Chen, Xingwu, et al.
Published: (2025)
by: Chen, Xingwu, et al.
Published: (2025)
On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Beyond First-Order: Training LLMs with Stochastic Conjugate Subgradients and AdamW
by: Zhang, Di, et al.
Published: (2025)
by: Zhang, Di, et al.
Published: (2025)
PAWN: Piece Value Analysis with Neural Networks
by: Tang, Ethan, et al.
Published: (2026)
by: Tang, Ethan, et al.
Published: (2026)
Understanding Adam Requires Better Rotation Dependent Assumptions
by: Zhang, Tianyue H., et al.
Published: (2024)
by: Zhang, Tianyue H., et al.
Published: (2024)
Stochastic Gradient Descent with Momentum is Algorithmically Stable
by: Lei, Yunwen, et al.
Published: (2026)
by: Lei, Yunwen, et al.
Published: (2026)
STGAN: Spatial-temporal Graph Autoregression Network for Pavement Distress Deterioration Prediction
by: Tong, Shilin, et al.
Published: (2025)
by: Tong, Shilin, et al.
Published: (2025)
Gradient Flow Convergence Guarantee for General Neural Network Architectures
by: Jakhmola, Yash
Published: (2025)
by: Jakhmola, Yash
Published: (2025)
Cross-Entropy Optimization for Hyperparameter Optimization in Stochastic Gradient-based Approaches to Train Deep Neural Networks
by: Li, Kevin, et al.
Published: (2024)
by: Li, Kevin, et al.
Published: (2024)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
by: Li, Yibo, et al.
Published: (2026)
by: Li, Yibo, et al.
Published: (2026)
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
by: Cloud, Alex, et al.
Published: (2024)
by: Cloud, Alex, et al.
Published: (2024)
Axiomatization of Gradient Smoothing in Neural Networks
by: Zhou, Linjiang, et al.
Published: (2024)
by: Zhou, Linjiang, et al.
Published: (2024)
SURGE: Surrogate Gradient Adaptation in Binary Neural Networks
by: Huang, Haoyu, et al.
Published: (2026)
by: Huang, Haoyu, et al.
Published: (2026)
Guiding Diffusion Models with Reinforcement Learning for Stable Molecule Generation
by: Zhou, Zhijian, et al.
Published: (2025)
by: Zhou, Zhijian, et al.
Published: (2025)
A Theoretical Understanding of Gradient Bias in Meta-Reinforcement Learning
by: Feng, Xidong, et al.
Published: (2021)
by: Feng, Xidong, et al.
Published: (2021)
Revisiting Adam for Streaming Reinforcement Learning
by: Gogianu, Florin, et al.
Published: (2026)
by: Gogianu, Florin, et al.
Published: (2026)
Gradient Inversion Attack on Graph Neural Networks
by: Sinha, Divya Anand, et al.
Published: (2024)
by: Sinha, Divya Anand, et al.
Published: (2024)
GradINN: Gradient Informed Neural Network
by: Aglietti, Filippo, et al.
Published: (2024)
by: Aglietti, Filippo, et al.
Published: (2024)
Gradient-Free Training of Quantized Neural Networks
by: Cohen, Noa, et al.
Published: (2024)
by: Cohen, Noa, et al.
Published: (2024)
Stochastic Weight Sharing for Bayesian Neural Networks
by: Lin, Moule, et al.
Published: (2025)
by: Lin, Moule, et al.
Published: (2025)
FedEGG: Federated Learning with Explicit Global Guidance
by: Zhai, Kun, et al.
Published: (2024)
by: Zhai, Kun, et al.
Published: (2024)
Norm-Hierarchy Transitions in Representation Learning: When and Why Neural Networks Abandon Shortcuts
by: Khanh, Truong Xuan, et al.
Published: (2026)
by: Khanh, Truong Xuan, et al.
Published: (2026)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
by: Jin, Ruinan, et al.
Published: (2026)
by: Jin, Ruinan, et al.
Published: (2026)
Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio
by: Wen, Ziqing, et al.
Published: (2026)
by: Wen, Ziqing, et al.
Published: (2026)
Understanding Pooling in Graph Neural Networks
by: Grattarola, Daniele, et al.
Published: (2021)
by: Grattarola, Daniele, et al.
Published: (2021)
Similar Items
-
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
by: Tang, Xuan, et al.
Published: (2025) -
Scaling Laws for Precision in High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2026) -
Learning under Quantization for High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2025) -
The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient
by: Li, Jichu, et al.
Published: (2026) -
The Implicit Bias of Adam on Separable Data
by: Zhang, Chenyang, et al.
Published: (2024)