Simple Convergence Proof of Adam From a Sign-like Descent Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, Hanyang, Qin, Shuang, Yu, Yue, Jiang, Fangqing, Wang, Hui, Lin, Zhouchen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
von: Peng, Hanyang, et al.
Veröffentlicht: (2025)
von: Peng, Hanyang, et al.
Veröffentlicht: (2025)
The Detection-Extraction Gap: Models Know the Answer Before They Can Say It
von: Wang, Hanyang, et al.
Veröffentlicht: (2026)
von: Wang, Hanyang, et al.
Veröffentlicht: (2026)
FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information
von: Hwang, Dongseong
Veröffentlicht: (2024)
von: Hwang, Dongseong
Veröffentlicht: (2024)
A Simple Model of Inference Scaling Laws
von: Levi, Noam
Veröffentlicht: (2024)
von: Levi, Noam
Veröffentlicht: (2024)
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
von: Chen, Zhirui, et al.
Veröffentlicht: (2024)
von: Chen, Zhirui, et al.
Veröffentlicht: (2024)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
von: Yang, Tong, et al.
Veröffentlicht: (2025)
von: Yang, Tong, et al.
Veröffentlicht: (2025)
Accelerating Convergence of Score-Based Diffusion Models, Provably
von: Li, Gen, et al.
Veröffentlicht: (2024)
von: Li, Gen, et al.
Veröffentlicht: (2024)
Why Does Stochastic Gradient Descent Slow Down in Low-Precision Training?
von: Yun, Vincent-Daniel
Veröffentlicht: (2025)
von: Yun, Vincent-Daniel
Veröffentlicht: (2025)
Fairness Overfitting in Machine Learning: An Information-Theoretic Perspective
von: Laakom, Firas, et al.
Veröffentlicht: (2025)
von: Laakom, Firas, et al.
Veröffentlicht: (2025)
FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large Models
von: Liu, Junkang, et al.
Veröffentlicht: (2025)
von: Liu, Junkang, et al.
Veröffentlicht: (2025)
Uncertainty Quantification and Data Efficiency in AI: An Information-Theoretic Perspective
von: Simeone, Osvaldo, et al.
Veröffentlicht: (2025)
von: Simeone, Osvaldo, et al.
Veröffentlicht: (2025)
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
von: Ouyang, Xu, et al.
Veröffentlicht: (2026)
von: Ouyang, Xu, et al.
Veröffentlicht: (2026)
Generalization Bounds: Perspectives from Information Theory and PAC-Bayes
von: Hellström, Fredrik, et al.
Veröffentlicht: (2023)
von: Hellström, Fredrik, et al.
Veröffentlicht: (2023)
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
von: Zhang, Yukun, et al.
Veröffentlicht: (2024)
von: Zhang, Yukun, et al.
Veröffentlicht: (2024)
A Deep Latent Factor Graph Clustering with Fairness-Utility Trade-off Perspective
von: Ghodsi, Siamak, et al.
Veröffentlicht: (2025)
von: Ghodsi, Siamak, et al.
Veröffentlicht: (2025)
The Blind Normalized Stein Variational Gradient Descent-Based Detection for Intelligent Random Access in Cellular IoT
von: Zhu, Xin, et al.
Veröffentlicht: (2024)
von: Zhu, Xin, et al.
Veröffentlicht: (2024)
Semi-supervised Batch Learning From Logged Data
von: Aminian, Gholamali, et al.
Veröffentlicht: (2022)
von: Aminian, Gholamali, et al.
Veröffentlicht: (2022)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
From MIM-Based GAN to Anomaly Detection:Event Probability Influence on Generative Adversarial Networks
von: She, Rui, et al.
Veröffentlicht: (2022)
von: She, Rui, et al.
Veröffentlicht: (2022)
WarpAdam: A new Adam optimizer based on Meta-Learning approach
von: Pan, Chengxi, et al.
Veröffentlicht: (2024)
von: Pan, Chengxi, et al.
Veröffentlicht: (2024)
Neural Channel Knowledge Map Assisted Scheduling Optimization of Active IRSs in Multi-User Systems
von: Chen, Xintong, et al.
Veröffentlicht: (2025)
von: Chen, Xintong, et al.
Veröffentlicht: (2025)
The Geometry of Knowing: From Possibilistic Ignorance to Probabilistic Certainty -- A Measure-Theoretic Framework for Epistemic Convergence
von: Jah, Moriba Kemessia
Veröffentlicht: (2026)
von: Jah, Moriba Kemessia
Veröffentlicht: (2026)
An Information Theoretic Perspective on Agentic System Design
von: He, Shizhe, et al.
Veröffentlicht: (2025)
von: He, Shizhe, et al.
Veröffentlicht: (2025)
Integrated Sensing-Communication-Computation for Edge Artificial Intelligence
von: Wen, Dingzhu, et al.
Veröffentlicht: (2023)
von: Wen, Dingzhu, et al.
Veröffentlicht: (2023)
Contrastive ECOC: Learning Output Codes for Adversarial Defense
von: Chou, Che-Yu, et al.
Veröffentlicht: (2025)
von: Chou, Che-Yu, et al.
Veröffentlicht: (2025)
In-Context Learning for MIMO Equalization Using Transformer-Based Sequence Models
von: Zecchin, Matteo, et al.
Veröffentlicht: (2023)
von: Zecchin, Matteo, et al.
Veröffentlicht: (2023)
GeoIB: Geometry-Aware Information Bottleneck via Statistical-Manifold Compression
von: Wang, Weiqi, et al.
Veröffentlicht: (2026)
von: Wang, Weiqi, et al.
Veröffentlicht: (2026)
Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
von: Wang, Qi, et al.
Veröffentlicht: (2025)
von: Wang, Qi, et al.
Veröffentlicht: (2025)
A Distance Measure for Random Permutation Set: From the Layer-2 Belief Structure Perspective
von: Cheng, Ruolan, et al.
Veröffentlicht: (2025)
von: Cheng, Ruolan, et al.
Veröffentlicht: (2025)
Estimating Mutual Information between Time Series and Temporal Event Sequences Across Diverse Analysis Tasks
von: Hu, Haoji, et al.
Veröffentlicht: (2026)
von: Hu, Haoji, et al.
Veröffentlicht: (2026)
Provable Privacy Advantages of Decentralized Federated Learning via Distributed Optimization
von: Yu, Wenrui, et al.
Veröffentlicht: (2024)
von: Yu, Wenrui, et al.
Veröffentlicht: (2024)
Relational Learning in Pre-Trained Models: A Theory from Hypergraph Recovery Perspective
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Conda: Column-Normalized Adam for Training Large Language Models Faster
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
Memorization-Compression Cycles Improve Generalization
von: Yu, Fangyuan
Veröffentlicht: (2025)
von: Yu, Fangyuan
Veröffentlicht: (2025)
CSRv2: Unlocking Ultra-Sparse Embeddings
von: Guo, Lixuan, et al.
Veröffentlicht: (2026)
von: Guo, Lixuan, et al.
Veröffentlicht: (2026)
Multi-Group Proportional Representation in Retrieval
von: Oesterling, Alex, et al.
Veröffentlicht: (2024)
von: Oesterling, Alex, et al.
Veröffentlicht: (2024)
Intent-Aware DRL-Based NOMA Uplink Dynamic Scheduler for IIoT
von: Mostafa, Salwa, et al.
Veröffentlicht: (2024)
von: Mostafa, Salwa, et al.
Veröffentlicht: (2024)
Compressing Chemistry Reveals Functional Groups
von: Sharma, Ruben, et al.
Veröffentlicht: (2025)
von: Sharma, Ruben, et al.
Veröffentlicht: (2025)
On the Limits of Self-Improving in Large Language Models: The Singularity Is Not Near Without Symbolic Model Synthesis
von: Zenil, Hector
Veröffentlicht: (2026)
von: Zenil, Hector
Veröffentlicht: (2026)
Deep Randomized Distributed Function Computation (DeepRDFC): Neural Distributed Channel Simulation
von: Bergström, Didrik, et al.
Veröffentlicht: (2026)
von: Bergström, Didrik, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
von: Peng, Hanyang, et al.
Veröffentlicht: (2025) -
The Detection-Extraction Gap: Models Know the Answer Before They Can Say It
von: Wang, Hanyang, et al.
Veröffentlicht: (2026) -
FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information
von: Hwang, Dongseong
Veröffentlicht: (2024) -
A Simple Model of Inference Scaling Laws
von: Levi, Noam
Veröffentlicht: (2024) -
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
von: Chen, Zhirui, et al.
Veröffentlicht: (2024)