Anon: Extrapolating Adaptivity Beyond SGD and Adam
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yiheng, Zhao, Kaiyan, Wu, Shaowu, Wang, Yiming, Wu, Jiajun, U, Leong Hou, Drew, Steve, Niu, Xiaoguang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HVAdam: A Full-Dimension Adaptive Optimizer
by: Zhang, Yiheng, et al.
Published: (2025)
by: Zhang, Yiheng, et al.
Published: (2025)
Efficient Diversity-based Experience Replay for Deep Reinforcement Learning
by: Zhao, Kaiyan, et al.
Published: (2024)
by: Zhao, Kaiyan, et al.
Published: (2024)
ANO: A Principled Approach to Robust Policy Optimization
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
Enhancing LLM Agents for Code Generation with Possibility and Pass-rate Prioritized Experience Replay
by: Chen, Yuyang, et al.
Published: (2024)
by: Chen, Yuyang, et al.
Published: (2024)
SPEAR: Soft Prompt Enhanced Anomaly Recognition for Time Series Data
by: Wei, Hanzhe, et al.
Published: (2025)
by: Wei, Hanzhe, et al.
Published: (2025)
Enhancing Equitable Access to AI in Housing and Homelessness System of Care through Federated Learning
by: Taib, Musa, et al.
Published: (2024)
by: Taib, Musa, et al.
Published: (2024)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
by: Peng, Hanyang, et al.
Published: (2025)
by: Peng, Hanyang, et al.
Published: (2025)
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024)
by: Zhu, Hanqing, et al.
Published: (2024)
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
by: Mukherjee, Sagnik, et al.
Published: (2026)
by: Mukherjee, Sagnik, et al.
Published: (2026)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
by: Jin, Ruinan, et al.
Published: (2026)
by: Jin, Ruinan, et al.
Published: (2026)
Structural Rationale Distillation via Reasoning Space Compression
by: Yang, Jialin, et al.
Published: (2026)
by: Yang, Jialin, et al.
Published: (2026)
VecCity: A Taxonomy-guided Library for Map Entity Representation Learning
by: Zhang, Wentao, et al.
Published: (2024)
by: Zhang, Wentao, et al.
Published: (2024)
Beyond First-Order: Training LLMs with Stochastic Conjugate Subgradients and AdamW
by: Zhang, Di, et al.
Published: (2025)
by: Zhang, Di, et al.
Published: (2025)
Beyond the False Trade-off: Adaptive EWC for Stealthy and Generalizable T2I Backdoors
by: Bowen, Lu, et al.
Published: (2026)
by: Bowen, Lu, et al.
Published: (2026)
Connectivity-Guided Sparsification of 2-FWL GNNs: Preserving Full Expressivity with Improved Efficiency
by: Chen, Rongqin, et al.
Published: (2025)
by: Chen, Rongqin, et al.
Published: (2025)
SAFLEX: Self-Adaptive Augmentation via Feature Label Extrapolation
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
Beyond Interpolation: Extrapolative Reasoning with Reinforcement Learning and Graph Neural Networks
by: Grillo, Niccolò, et al.
Published: (2025)
by: Grillo, Niccolò, et al.
Published: (2025)
Multi-Value Alignment for LLMs via Value Decorrelation and Extrapolation
by: Xu, Hefei, et al.
Published: (2025)
by: Xu, Hefei, et al.
Published: (2025)
REE-TTT: Highly Adaptive Radar Echo Extrapolation Based on Test-Time Training
by: Di, Xin, et al.
Published: (2026)
by: Di, Xin, et al.
Published: (2026)
Enhancing DP-SGD through Non-monotonous Adaptive Scaling Gradient Weight
by: Huang, Tao, et al.
Published: (2024)
by: Huang, Tao, et al.
Published: (2024)
Transferable Delay-Aware Reinforcement Learning via Implicit Causal Graph Modeling
by: Zhao, Chenran, et al.
Published: (2026)
by: Zhao, Chenran, et al.
Published: (2026)
Accumulative SGD Influence Estimation for Data Attribution
by: Shi, Yunxiao, et al.
Published: (2025)
by: Shi, Yunxiao, et al.
Published: (2025)
Representations learnt by SGD and Adaptive learning rules: Conditions that vary sparsity and selectivity in neural networks
by: Park, Jin Hyun
Published: (2022)
by: Park, Jin Hyun
Published: (2022)
Mesa-Extrapolation: A Weave Position Encoding Method for Enhanced Extrapolation in LLMs
by: Ma, Xin, et al.
Published: (2024)
by: Ma, Xin, et al.
Published: (2024)
Adam-mini: Use Fewer Learning Rates To Gain More
by: Zhang, Yushun, et al.
Published: (2024)
by: Zhang, Yushun, et al.
Published: (2024)
Conda: Column-Normalized Adam for Training Large Language Models Faster
by: Wang, Junjie, et al.
Published: (2025)
by: Wang, Junjie, et al.
Published: (2025)
Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents
by: Li, Ziming, et al.
Published: (2026)
by: Li, Ziming, et al.
Published: (2026)
Bootstrap SGD: Algorithmic Stability and Robustness
by: Christmann, Andreas, et al.
Published: (2024)
by: Christmann, Andreas, et al.
Published: (2024)
Beyond the Kolmogorov Barrier: A Learnable Weighted Hybrid Autoencoder for Model Order Reduction
by: Somasekharan, Nithin, et al.
Published: (2024)
by: Somasekharan, Nithin, et al.
Published: (2024)
The Extrapolation Power of Implicit Models
by: Decugis, Juliette, et al.
Published: (2024)
by: Decugis, Juliette, et al.
Published: (2024)
Identifying Representations for Intervention Extrapolation
by: Saengkyongam, Sorawit, et al.
Published: (2023)
by: Saengkyongam, Sorawit, et al.
Published: (2023)
AGMixup: Adaptive Graph Mixup for Semi-supervised Node Classification
by: Lu, Weigang, et al.
Published: (2024)
by: Lu, Weigang, et al.
Published: (2024)
Towards Understanding Extrapolation: a Causal Lens
by: Kong, Lingjing, et al.
Published: (2025)
by: Kong, Lingjing, et al.
Published: (2025)
In-Run Data Shapley for Adam Optimizer
by: Ding, Meng, et al.
Published: (2026)
by: Ding, Meng, et al.
Published: (2026)
WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales
by: Prinster, Drew, et al.
Published: (2025)
by: Prinster, Drew, et al.
Published: (2025)
SOAP: Improving and Stabilizing Shampoo using Adam
by: Vyas, Nikhil, et al.
Published: (2024)
by: Vyas, Nikhil, et al.
Published: (2024)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
DynaSwarm: Dynamically Graph Structure Selection for LLM-based Multi-agent System
by: Leong, Hui Yi, et al.
Published: (2025)
by: Leong, Hui Yi, et al.
Published: (2025)
Gradient Extrapolation-Based Policy Optimization
by: Swapnil, Ismam Nur, et al.
Published: (2026)
by: Swapnil, Ismam Nur, et al.
Published: (2026)
DP-Adam-AC: Privacy-preserving Fine-Tuning of Localizable Language Models Using Adam Optimization with Adaptive Clipping
by: Yang, Ruoxing
Published: (2025)
by: Yang, Ruoxing
Published: (2025)
Similar Items
-
HVAdam: A Full-Dimension Adaptive Optimizer
by: Zhang, Yiheng, et al.
Published: (2025) -
Efficient Diversity-based Experience Replay for Deep Reinforcement Learning
by: Zhao, Kaiyan, et al.
Published: (2024) -
ANO: A Principled Approach to Robust Policy Optimization
by: Zhang, Yiheng, et al.
Published: (2026) -
Enhancing LLM Agents for Code Generation with Possibility and Pass-rate Prioritized Experience Replay
by: Chen, Yuyang, et al.
Published: (2024) -
SPEAR: Soft Prompt Enhanced Anomaly Recognition for Time Series Data
by: Wei, Hanzhe, et al.
Published: (2025)