Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Huangyu, Yang, Jingqin, Xu, Qianqian, Teng, Jiaye |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reparameterization Flow Policy Optimization
by: Zhong, Hai, et al.
Published: (2026)
by: Zhong, Hai, et al.
Published: (2026)
Reparameterization Proximal Policy Optimization
by: Zhong, Hai, et al.
Published: (2025)
by: Zhong, Hai, et al.
Published: (2025)
Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap
by: Sun, Yifan, et al.
Published: (2025)
by: Sun, Yifan, et al.
Published: (2025)
The Geometry of Multi-Task Grokking: Transverse Instability, Superposition, and Weight Decay Phase Structure
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
What Makes Looped Transformers Perform Better Than Non-Recursive Ones
by: Gong, Zixuan, et al.
Published: (2025)
by: Gong, Zixuan, et al.
Published: (2025)
From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight
by: Fu, Xiaoliang, et al.
Published: (2026)
by: Fu, Xiaoliang, et al.
Published: (2026)
Adaptive Weighted Loss for Sequential Recommendations on Sparse Domains
by: Mittal, Akshay, et al.
Published: (2025)
by: Mittal, Akshay, et al.
Published: (2025)
Provably Invincible Adversarial Attacks on Reinforcement Learning Systems: A Rate-Distortion Information-Theoretic Approach
by: Lu, Ziqing, et al.
Published: (2025)
by: Lu, Ziqing, et al.
Published: (2025)
On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm Perspective
by: Xie, Zeke, et al.
Published: (2020)
by: Xie, Zeke, et al.
Published: (2020)
Predictive Inference With Fast Feature Conformal Prediction
by: Tang, Zihao, et al.
Published: (2024)
by: Tang, Zihao, et al.
Published: (2024)
Optimal Linear Decay Learning Rate Schedules and Further Refinements
by: Defazio, Aaron, et al.
Published: (2023)
by: Defazio, Aaron, et al.
Published: (2023)
$\clubsuit$ CLOVER $\clubsuit$: Probabilistic Forecasting with Coherent Learning Objective Reparameterization
by: Olivares, Kin G., et al.
Published: (2023)
by: Olivares, Kin G., et al.
Published: (2023)
SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
by: Zheng, Zhi, et al.
Published: (2025)
by: Zheng, Zhi, et al.
Published: (2025)
Beyond Alignment: Expanding Reasoning Capacity via Manifold-Reshaping Policy Optimization
by: Wang, Dayu, et al.
Published: (2026)
by: Wang, Dayu, et al.
Published: (2026)
kNN-Graph: An adaptive graph model for $k$-nearest neighbors
by: Li, Jiaye, et al.
Published: (2026)
by: Li, Jiaye, et al.
Published: (2026)
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
by: Dremov, Aleksandr, et al.
Published: (2025)
by: Dremov, Aleksandr, et al.
Published: (2025)
KTO: Model Alignment as Prospect Theoretic Optimization
by: Ethayarajh, Kawin, et al.
Published: (2024)
by: Ethayarajh, Kawin, et al.
Published: (2024)
Meta Additive Model: Interpretable Sparse Learning With Auto Weighting
by: Zhang, Xuelin, et al.
Published: (2026)
by: Zhang, Xuelin, et al.
Published: (2026)
A Theoretical Analysis of Efficiency Constrained Utility-Privacy Bi-Objective Optimization in Federated Learning
by: Gu, Hanlin, et al.
Published: (2023)
by: Gu, Hanlin, et al.
Published: (2023)
On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery
by: Liu, Renpu, et al.
Published: (2024)
by: Liu, Renpu, et al.
Published: (2024)
Enhancing In-Context Learning Performance with just SVD-Based Weight Pruning: A Theoretical Perspective
by: Yao, Xinhao, et al.
Published: (2024)
by: Yao, Xinhao, et al.
Published: (2024)
Optimizing Partial Area Under the Top-k Curve: Theory and Practice
by: Wang, Zitai, et al.
Published: (2022)
by: Wang, Zitai, et al.
Published: (2022)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
by: Fan, Zhiyuan, et al.
Published: (2025)
by: Fan, Zhiyuan, et al.
Published: (2025)
Enhancing DP-SGD through Non-monotonous Adaptive Scaling Gradient Weight
by: Huang, Tao, et al.
Published: (2024)
by: Huang, Tao, et al.
Published: (2024)
Contrastive Learning Is Spectral Clustering On Similarity Graph
by: Tan, Zhiquan, et al.
Published: (2023)
by: Tan, Zhiquan, et al.
Published: (2023)
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
by: Gong, Zixuan, et al.
Published: (2025)
by: Gong, Zixuan, et al.
Published: (2025)
Adaptive Memory Decay for Log-Linear Attention
by: Amin, Yaxita, et al.
Published: (2026)
by: Amin, Yaxita, et al.
Published: (2026)
ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization
by: Yang, Letian, et al.
Published: (2026)
by: Yang, Letian, et al.
Published: (2026)
Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks
by: Ma, Wenquan, et al.
Published: (2026)
by: Ma, Wenquan, et al.
Published: (2026)
CAdam: Confidence-Based Optimization for Online Learning
by: Wang, Shaowen, et al.
Published: (2024)
by: Wang, Shaowen, et al.
Published: (2024)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
by: He, Di, et al.
Published: (2025)
by: He, Di, et al.
Published: (2025)
HVAdam: A Full-Dimension Adaptive Optimizer
by: Zhang, Yiheng, et al.
Published: (2025)
by: Zhang, Yiheng, et al.
Published: (2025)
FeDeRA:Efficient Fine-tuning of Language Models in Federated Learning Leveraging Weight Decomposition
by: Yan, Yuxuan, et al.
Published: (2024)
by: Yan, Yuxuan, et al.
Published: (2024)
No More Adam: Learning Rate Scaling at Initialization is All You Need
by: Xu, Minghao, et al.
Published: (2024)
by: Xu, Minghao, et al.
Published: (2024)
Adaptive Correlation-Weighted Intrinsic Rewards for Reinforcement Learning
by: Nguyen, Viet Bac, et al.
Published: (2026)
by: Nguyen, Viet Bac, et al.
Published: (2026)
Reinforcement Learning With Sparse-Executing Actions via Sparsity Regularization
by: Pang, Jing-Cheng, et al.
Published: (2021)
by: Pang, Jing-Cheng, et al.
Published: (2021)
Reparameterized Tensor Ring Functional Decomposition for Multi-Dimensional Data Recovery
by: Xu, Yangyang, et al.
Published: (2026)
by: Xu, Yangyang, et al.
Published: (2026)
Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
by: Kang, Hyeongyu, et al.
Published: (2025)
by: Kang, Hyeongyu, et al.
Published: (2025)
EDGE: A Theoretical Framework for Misconception-Aware Adaptive Learning
by: Verma, Ananda Prakash
Published: (2025)
by: Verma, Ananda Prakash
Published: (2025)
Hindsight-Guided Momentum (HGM) Optimizer: An Approach to Adaptive Learning Rate
by: Sarkar, Krisanu
Published: (2025)
by: Sarkar, Krisanu
Published: (2025)
Similar Items
-
Reparameterization Flow Policy Optimization
by: Zhong, Hai, et al.
Published: (2026) -
Reparameterization Proximal Policy Optimization
by: Zhong, Hai, et al.
Published: (2025) -
Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap
by: Sun, Yifan, et al.
Published: (2025) -
The Geometry of Multi-Task Grokking: Transverse Instability, Superposition, and Weight Decay Phase Structure
by: Xu, Yongzhong
Published: (2026) -
What Makes Looped Transformers Perform Better Than Non-Recursive Ones
by: Gong, Zixuan, et al.
Published: (2025)