Gespeichert in:
| Hauptverfasser: | Nguyen, Son, Liu, Bo, Chen, Lizhang, Liu, Qiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.07488 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Memory-Efficient Optimization with Factorized Hamiltonian Descent
von: Nguyen, Son, et al.
Veröffentlicht: (2024)
von: Nguyen, Son, et al.
Veröffentlicht: (2024)
Cautious Optimizers: Improving Training with One Line of Code
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts
von: Chen, Lizhang, et al.
Veröffentlicht: (2023)
von: Chen, Lizhang, et al.
Veröffentlicht: (2023)
Muon Optimizes Under Spectral Norm Constraints
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
Memory-Efficient LLM Training with Online Subspace Descent
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
DeMo: Decoupled Momentum Optimization
von: Peng, Bowen, et al.
Veröffentlicht: (2024)
von: Peng, Bowen, et al.
Veröffentlicht: (2024)
Training-Free Looped Transformers
von: Chen, Lizhang, et al.
Veröffentlicht: (2026)
von: Chen, Lizhang, et al.
Veröffentlicht: (2026)
Structured Preconditioners in Adaptive Optimization: A Unified Analysis
von: Xie, Shuo, et al.
Veröffentlicht: (2025)
von: Xie, Shuo, et al.
Veröffentlicht: (2025)
Adaptive Preconditioners Trigger Loss Spikes in Adam
von: Bai, Zhiwei, et al.
Veröffentlicht: (2025)
von: Bai, Zhiwei, et al.
Veröffentlicht: (2025)
Momentum Guidance: Plug-and-Play Guidance for Flow Models
von: Liao, Runlong, et al.
Veröffentlicht: (2026)
von: Liao, Runlong, et al.
Veröffentlicht: (2026)
Communication Efficient Distributed Training with Distributed Lion
von: Liu, Bo, et al.
Veröffentlicht: (2024)
von: Liu, Bo, et al.
Veröffentlicht: (2024)
Taming Preconditioner Drift: Unlocking the Potential of Second-Order Optimizers for Federated Learning on Non-IID Data
von: Liu, Junkang, et al.
Veröffentlicht: (2026)
von: Liu, Junkang, et al.
Veröffentlicht: (2026)
$ϕ$-Balancing for Mixture-of-Experts Training
von: Chen, Lizhang, et al.
Veröffentlicht: (2026)
von: Chen, Lizhang, et al.
Veröffentlicht: (2026)
AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies
von: Hu, Xixi, et al.
Veröffentlicht: (2024)
von: Hu, Xixi, et al.
Veröffentlicht: (2024)
Cautious Weight Decay
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
SAMix: Calibrated and Accurate Continual Learning via Sphere-Adaptive Mixup and Neural Collapse
von: Dang, Trung-Anh, et al.
Veröffentlicht: (2025)
von: Dang, Trung-Anh, et al.
Veröffentlicht: (2025)
Curvature-Informed SGD via General Purpose Lie-Group Preconditioners
von: Pooladzandi, Omead, et al.
Veröffentlicht: (2024)
von: Pooladzandi, Omead, et al.
Veröffentlicht: (2024)
CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM
von: Nguyen, Son, et al.
Veröffentlicht: (2026)
von: Nguyen, Son, et al.
Veröffentlicht: (2026)
Diagonally-Weighted Generalized Method of Moments Estimation for Gaussian Mixture Modeling
von: Zhang, Liu, et al.
Veröffentlicht: (2025)
von: Zhang, Liu, et al.
Veröffentlicht: (2025)
Graph Neural Preconditioners for Iterative Solutions of Sparse Linear Systems
von: Chen, Jie
Veröffentlicht: (2024)
von: Chen, Jie
Veröffentlicht: (2024)
Learning Sparse Approximate Inverse Preconditioners for Conjugate Gradient Solvers on GPUs
von: Yang, Zherui, et al.
Veröffentlicht: (2025)
von: Yang, Zherui, et al.
Veröffentlicht: (2025)
A New Perspective on Shampoo's Preconditioner
von: Morwani, Depen, et al.
Veröffentlicht: (2024)
von: Morwani, Depen, et al.
Veröffentlicht: (2024)
Gaussian Processes Sampling with Sparse Grids under Additive Schwarz Preconditioner
von: Chen, Haoyuan, et al.
Veröffentlicht: (2024)
von: Chen, Haoyuan, et al.
Veröffentlicht: (2024)
Improving Probabilistic Diffusion Models With Optimal Diagonal Covariance Matching
von: Ou, Zijing, et al.
Veröffentlicht: (2024)
von: Ou, Zijing, et al.
Veröffentlicht: (2024)
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
von: Zhang, Minxin, et al.
Veröffentlicht: (2026)
von: Zhang, Minxin, et al.
Veröffentlicht: (2026)
Regime-Adaptive Bayesian Optimization via Dirichlet Process Mixtures of Gaussian Processes
von: Zhang, Yan, et al.
Veröffentlicht: (2026)
von: Zhang, Yan, et al.
Veröffentlicht: (2026)
A Self-Attentive Meta-Optimizer with Group-Adaptive Learning Rates and Weight Decay
von: Zhao, JiangBo, et al.
Veröffentlicht: (2026)
von: Zhao, JiangBo, et al.
Veröffentlicht: (2026)
An Experimental Study of Semantic Continuity for Deep Learning Models
von: Wu, Shangxi, et al.
Veröffentlicht: (2020)
von: Wu, Shangxi, et al.
Veröffentlicht: (2020)
Adaptive Estimation and Inference in Conditional Moment Models via the Discrepancy Principle
von: Tan, Jiyuan, et al.
Veröffentlicht: (2026)
von: Tan, Jiyuan, et al.
Veröffentlicht: (2026)
Spectral Embeddings Leak Graph Topology: Theory, Benchmark, and Adaptive Reconstruction
von: Nguyen-Cong, Thinh, et al.
Veröffentlicht: (2026)
von: Nguyen-Cong, Thinh, et al.
Veröffentlicht: (2026)
Diagonal Adaptive Non-local Observables on Quantum Neural Networks
von: Tseng, Huan-Hsin, et al.
Veröffentlicht: (2026)
von: Tseng, Huan-Hsin, et al.
Veröffentlicht: (2026)
Optimization Insights into Deep Diagonal Linear Networks
von: Labarrière, Hippolyte, et al.
Veröffentlicht: (2024)
von: Labarrière, Hippolyte, et al.
Veröffentlicht: (2024)
Generative modeling of Sparse Approximate Inverse Preconditioners
von: Li, Mou, et al.
Veröffentlicht: (2024)
von: Li, Mou, et al.
Veröffentlicht: (2024)
Diagonal Over-parameterization in Reproducing Kernel Hilbert Spaces as an Adaptive Feature Model: Generalization and Adaptivity
von: Li, Yicheng, et al.
Veröffentlicht: (2025)
von: Li, Yicheng, et al.
Veröffentlicht: (2025)
Preconditioners for the Stochastic Training of Neural Fields
von: Chng, Shin-Fang, et al.
Veröffentlicht: (2024)
von: Chng, Shin-Fang, et al.
Veröffentlicht: (2024)
Improving Rectified Flow with Boundary Conditions
von: Hu, Xixi, et al.
Veröffentlicht: (2025)
von: Hu, Xixi, et al.
Veröffentlicht: (2025)
Improving Deep Knowledge Tracing via Gated Architectures and Adaptive Optimization
von: Shukurlu, Altun
Veröffentlicht: (2025)
von: Shukurlu, Altun
Veröffentlicht: (2025)
Adaptive Moment Estimation Optimization Algorithm Using Projection Gradient for Deep Learning
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
GADPN: Graph Adaptive Denoising and Perturbation Networks via Singular Value Decomposition
von: Deng, Hao, et al.
Veröffentlicht: (2026)
von: Deng, Hao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Memory-Efficient Optimization with Factorized Hamiltonian Descent
von: Nguyen, Son, et al.
Veröffentlicht: (2024) -
Cautious Optimizers: Improving Training with One Line of Code
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024) -
Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts
von: Chen, Lizhang, et al.
Veröffentlicht: (2023) -
Muon Optimizes Under Spectral Norm Constraints
von: Chen, Lizhang, et al.
Veröffentlicht: (2025) -
Memory-Efficient LLM Training with Online Subspace Descent
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)