A Self-Attentive Meta-Optimizer with Group-Adaptive Learning Rates and Weight Decay
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, JiangBo, Liu, ZhaoXin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
von: Xu, Huangyu, et al.
Veröffentlicht: (2026)
von: Xu, Huangyu, et al.
Veröffentlicht: (2026)
Deep Learning Optimization Using Self-Adaptive Weighted Auxiliary Variables
von: Liu, Yaru, et al.
Veröffentlicht: (2025)
von: Liu, Yaru, et al.
Veröffentlicht: (2025)
Learning to Forget: Continual Learning with Adaptive Weight Decay
von: Ramesh, Aditya A., et al.
Veröffentlicht: (2026)
von: Ramesh, Aditya A., et al.
Veröffentlicht: (2026)
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
von: Kosson, Atli, et al.
Veröffentlicht: (2025)
von: Kosson, Atli, et al.
Veröffentlicht: (2025)
Optimal Learning-Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay
von: Li, Binghui, et al.
Veröffentlicht: (2026)
von: Li, Binghui, et al.
Veröffentlicht: (2026)
Deep Learning meets Nonparametric Regression: Are Weight-Decayed DNNs Locally Adaptive?
von: Zhang, Kaiqi, et al.
Veröffentlicht: (2022)
von: Zhang, Kaiqi, et al.
Veröffentlicht: (2022)
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2025)
FedAWA: Adaptive Optimization of Aggregation Weights in Federated Learning Using Client Vectors
von: Shi, Changlong, et al.
Veröffentlicht: (2025)
von: Shi, Changlong, et al.
Veröffentlicht: (2025)
Tune without Validation: Searching for Learning Rate and Weight Decay on Training Sets
von: Brigato, Lorenzo, et al.
Veröffentlicht: (2024)
von: Brigato, Lorenzo, et al.
Veröffentlicht: (2024)
On the Role of Weight Decay in Collaborative Filtering: A Popularity Perspective
von: Loveland, Donald, et al.
Veröffentlicht: (2025)
von: Loveland, Donald, et al.
Veröffentlicht: (2025)
Cautious Weight Decay
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
Correction of Decoupled Weight Decay
von: Chou, Jason Chuan-Chih
Veröffentlicht: (2025)
von: Chou, Jason Chuan-Chih
Veröffentlicht: (2025)
Understanding the Generalization Benefits of Late Learning Rate Decay
von: Ren, Yinuo, et al.
Veröffentlicht: (2024)
von: Ren, Yinuo, et al.
Veröffentlicht: (2024)
OUIDecay: Adaptive Layer-wise Weight Decay for CNNs Using Online Activation Patterns
von: Fernández-Hernández, Alberto, et al.
Veröffentlicht: (2026)
von: Fernández-Hernández, Alberto, et al.
Veröffentlicht: (2026)
FedLWS: Federated Learning with Adaptive Layer-wise Weight Shrinking
von: Shi, Changlong, et al.
Veröffentlicht: (2025)
von: Shi, Changlong, et al.
Veröffentlicht: (2025)
Heterogeneous Graph Contrastive Learning with Meta-path Contexts and Adaptively Weighted Negative Samples
von: Yu, Jianxiang, et al.
Veröffentlicht: (2022)
von: Yu, Jianxiang, et al.
Veröffentlicht: (2022)
Taming LLMs by Scaling Learning Rates with Gradient Grouping
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
The Lifecycle of the Spectral Edge: From Gradient Learning to Weight-Decay Compression
von: Xu, Yongzhong
Veröffentlicht: (2026)
von: Xu, Yongzhong
Veröffentlicht: (2026)
Why Do We Need Weight Decay in Modern Deep Learning?
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2023)
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2023)
Temperature as a Meta-Policy: Adaptive Temperature in LLM Reinforcement Learning
von: Dang, Haoran, et al.
Veröffentlicht: (2026)
von: Dang, Haoran, et al.
Veröffentlicht: (2026)
Task-free Adaptive Meta Black-box Optimization
von: Wang, Chao, et al.
Veröffentlicht: (2026)
von: Wang, Chao, et al.
Veröffentlicht: (2026)
Weighted Temporal Decay Loss for Learning Wearable PPG Data with Sparse Clinical Labels
von: Chung, Yunsung, et al.
Veröffentlicht: (2026)
von: Chung, Yunsung, et al.
Veröffentlicht: (2026)
Action-Attentive Deep Reinforcement Learning for Autonomous Alignment of Beamlines
von: Wang, Siyu, et al.
Veröffentlicht: (2024)
von: Wang, Siyu, et al.
Veröffentlicht: (2024)
Urban Region Representation Learning with Attentive Fusion
von: Sun, Fengze, et al.
Veröffentlicht: (2023)
von: Sun, Fengze, et al.
Veröffentlicht: (2023)
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
von: Kosson, Atli, et al.
Veröffentlicht: (2023)
von: Kosson, Atli, et al.
Veröffentlicht: (2023)
The Rank and Gradient Lost in Non-stationarity: Sample Weight Decay for Mitigating Plasticity Loss in Reinforcement Learning
von: Wu, Zihao, et al.
Veröffentlicht: (2026)
von: Wu, Zihao, et al.
Veröffentlicht: (2026)
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
von: Song, Guanghui, et al.
Veröffentlicht: (2025)
von: Song, Guanghui, et al.
Veröffentlicht: (2025)
Optimal Linear Decay Learning Rate Schedules and Further Refinements
von: Defazio, Aaron, et al.
Veröffentlicht: (2023)
von: Defazio, Aaron, et al.
Veröffentlicht: (2023)
Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
von: Lu, Yining, et al.
Veröffentlicht: (2025)
von: Lu, Yining, et al.
Veröffentlicht: (2025)
Group-in-Group Policy Optimization for LLM Agent Training
von: Feng, Lang, et al.
Veröffentlicht: (2025)
von: Feng, Lang, et al.
Veröffentlicht: (2025)
An Energy-Based Self-Adaptive Learning Rate for Stochastic Gradient Descent: Enhancing Unconstrained Optimization with VAV method
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
AdaLomo: Low-memory Optimization with Adaptive Learning Rate
von: Lv, Kai, et al.
Veröffentlicht: (2023)
von: Lv, Kai, et al.
Veröffentlicht: (2023)
Noise-Adaptive Layerwise Learning Rates: Accelerating Geometry-Aware Optimization for Deep Neural Network Training
von: Hao, Jie, et al.
Veröffentlicht: (2025)
von: Hao, Jie, et al.
Veröffentlicht: (2025)
Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
von: He, Di, et al.
Veröffentlicht: (2025)
von: He, Di, et al.
Veröffentlicht: (2025)
Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning
von: Liu, Zehao, et al.
Veröffentlicht: (2026)
von: Liu, Zehao, et al.
Veröffentlicht: (2026)
Adaptive Dual-Weighting Framework for Federated Learning via Out-of-Distribution Detection
von: Ling, Zhiwei, et al.
Veröffentlicht: (2026)
von: Ling, Zhiwei, et al.
Veröffentlicht: (2026)
Improving Learning to Optimize Using Parameter Symmetries
von: Zamir, Guy, et al.
Veröffentlicht: (2025)
von: Zamir, Guy, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
von: Xu, Huangyu, et al.
Veröffentlicht: (2026) -
Deep Learning Optimization Using Self-Adaptive Weighted Auxiliary Variables
von: Liu, Yaru, et al.
Veröffentlicht: (2025) -
Learning to Forget: Continual Learning with Adaptive Weight Decay
von: Ramesh, Aditya A., et al.
Veröffentlicht: (2026) -
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
von: Kosson, Atli, et al.
Veröffentlicht: (2025) -
Optimal Learning-Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay
von: Li, Binghui, et al.
Veröffentlicht: (2026)