Structured Preconditioners in Adaptive Optimization: A Unified Analysis
Fuente:
arXiv
Guardado en:
| Autores principales: | Xie, Shuo, Wang, Tianhao, Reddi, Sashank, Kumar, Sanjiv, Li, Zhiyuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reasoning with Latent Thoughts: On the Power of Looped Transformers
por: Saunshi, Nikunj, et al.
Publicado: (2025)
por: Saunshi, Nikunj, et al.
Publicado: (2025)
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
por: Xie, Shuo, et al.
Publicado: (2025)
por: Xie, Shuo, et al.
Publicado: (2025)
Landscape-Aware Growing: The Power of a Little LAG
por: Karp, Stefani, et al.
Publicado: (2024)
por: Karp, Stefani, et al.
Publicado: (2024)
On the Role of Depth and Looping for In-Context Learning with Task Diversity
por: Gatmiry, Khashayar, et al.
Publicado: (2024)
por: Gatmiry, Khashayar, et al.
Publicado: (2024)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
por: Gatmiry, Khashayar, et al.
Publicado: (2024)
por: Gatmiry, Khashayar, et al.
Publicado: (2024)
Simplicity Bias via Global Convergence of Sharpness Minimization
por: Gatmiry, Khashayar, et al.
Publicado: (2024)
por: Gatmiry, Khashayar, et al.
Publicado: (2024)
On the Inductive Bias of Stacking Towards Improving Reasoning
por: Saunshi, Nikunj, et al.
Publicado: (2024)
por: Saunshi, Nikunj, et al.
Publicado: (2024)
Efficient Stagewise Pretraining via Progressive Subnetworks
por: Panigrahi, Abhishek, et al.
Publicado: (2024)
por: Panigrahi, Abhishek, et al.
Publicado: (2024)
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
por: Yadav, Robin, et al.
Publicado: (2025)
por: Yadav, Robin, et al.
Publicado: (2025)
Improving Adaptive Moment Optimization via Preconditioner Diagonalization
por: Nguyen, Son, et al.
Publicado: (2025)
por: Nguyen, Son, et al.
Publicado: (2025)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
por: Xie, Shuo, et al.
Publicado: (2024)
por: Xie, Shuo, et al.
Publicado: (2024)
Adaptive Preconditioners Trigger Loss Spikes in Adam
por: Bai, Zhiwei, et al.
Publicado: (2025)
por: Bai, Zhiwei, et al.
Publicado: (2025)
Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity
por: Xie, Shuo, et al.
Publicado: (2024)
por: Xie, Shuo, et al.
Publicado: (2024)
Efficient Document Ranking with Learnable Late Interactions
por: Ji, Ziwei, et al.
Publicado: (2024)
por: Ji, Ziwei, et al.
Publicado: (2024)
Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation
por: Lukasik, Michal, et al.
Publicado: (2025)
por: Lukasik, Michal, et al.
Publicado: (2025)
Deep Learning Agents Trained For Avoidance Behave Like Hawks And Doves
por: Reddi, Aryaman
Publicado: (2025)
por: Reddi, Aryaman
Publicado: (2025)
Unifying Sequences, Structures, and Descriptions for Any-to-Any Protein Generation with the Large Multimodal Model HelixProtX
por: Chen, Zhiyuan, et al.
Publicado: (2024)
por: Chen, Zhiyuan, et al.
Publicado: (2024)
Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
por: Mohamadi, Mohamad Amin, et al.
Publicado: (2025)
por: Mohamadi, Mohamad Amin, et al.
Publicado: (2025)
TinyTorch: Building Machine Learning Systems from First Principles
por: Reddi, Vijay Janapa
Publicado: (2026)
por: Reddi, Vijay Janapa
Publicado: (2026)
Generative modeling of Sparse Approximate Inverse Preconditioners
por: Li, Mou, et al.
Publicado: (2024)
por: Li, Mou, et al.
Publicado: (2024)
A Non-asymptotic Analysis for Learning and Applying a Preconditioner in MCMC
por: Hird, Max, et al.
Publicado: (2026)
por: Hird, Max, et al.
Publicado: (2026)
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
por: Rawat, Ankit Singh, et al.
Publicado: (2024)
por: Rawat, Ankit Singh, et al.
Publicado: (2024)
Reasoning Distillation for Lightweight Automated Program Repair
por: Balasubramanian, Aanand, et al.
Publicado: (2026)
por: Balasubramanian, Aanand, et al.
Publicado: (2026)
A New Perspective on Shampoo's Preconditioner
por: Morwani, Depen, et al.
Publicado: (2024)
por: Morwani, Depen, et al.
Publicado: (2024)
Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
por: Zhai, Zhiyuan, et al.
Publicado: (2026)
por: Zhai, Zhiyuan, et al.
Publicado: (2026)
Are More Tokens Rational? Inference-Time Scaling in Language Models as Adaptive Resource Rationality
por: Hu, Zhimin, et al.
Publicado: (2026)
por: Hu, Zhimin, et al.
Publicado: (2026)
Investigation of Compressor Cascade Flow Using Physics- Informed Neural Networks with Adaptive Learning Strategy
por: Li, Zhihui, et al.
Publicado: (2023)
por: Li, Zhihui, et al.
Publicado: (2023)
Unified Convergence Analysis for Adaptive Optimization with Moving Average Estimator
por: Guo, Zhishuai, et al.
Publicado: (2021)
por: Guo, Zhishuai, et al.
Publicado: (2021)
Curvature-Informed SGD via General Purpose Lie-Group Preconditioners
por: Pooladzandi, Omead, et al.
Publicado: (2024)
por: Pooladzandi, Omead, et al.
Publicado: (2024)
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
por: Kwon, Soo Min, et al.
Publicado: (2026)
por: Kwon, Soo Min, et al.
Publicado: (2026)
Unified Transfer Learning Models in High-Dimensional Linear Regression
por: Liu, Shuo Shuo
Publicado: (2023)
por: Liu, Shuo Shuo
Publicado: (2023)
Enhanced High-Dimensional Data Visualization through Adaptive Multi-Scale Manifold Embedding
por: Ni, Tianhao, et al.
Publicado: (2025)
por: Ni, Tianhao, et al.
Publicado: (2025)
Understanding the Countably Infinite: Neural Network Models of the Successor Function and its Acquisition
por: Gupta, Vima, et al.
Publicado: (2023)
por: Gupta, Vima, et al.
Publicado: (2023)
The Marginal Value of Momentum for Small Learning Rate SGD
por: Wang, Runzhe, et al.
Publicado: (2023)
por: Wang, Runzhe, et al.
Publicado: (2023)
Optimizing Few-Step Generation with Adaptive Matching Distillation
por: Bai, Lichen, et al.
Publicado: (2026)
por: Bai, Lichen, et al.
Publicado: (2026)
Generative AI Agents in Autonomous Machines: A Safety Perspective
por: Jabbour, Jason, et al.
Publicado: (2024)
por: Jabbour, Jason, et al.
Publicado: (2024)
Relative Kinetic Utility for Reasoning-Aware Structural Pruning in Large Language Models
por: Qian, Tianhao
Publicado: (2026)
por: Qian, Tianhao
Publicado: (2026)
Learning Sparse Approximate Inverse Preconditioners for Conjugate Gradient Solvers on GPUs
por: Yang, Zherui, et al.
Publicado: (2025)
por: Yang, Zherui, et al.
Publicado: (2025)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
por: Kovalev, Dmitry, et al.
Publicado: (2025)
por: Kovalev, Dmitry, et al.
Publicado: (2025)
FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning
por: Li, Jiaoyang, et al.
Publicado: (2025)
por: Li, Jiaoyang, et al.
Publicado: (2025)
Ejemplares similares
-
Reasoning with Latent Thoughts: On the Power of Looped Transformers
por: Saunshi, Nikunj, et al.
Publicado: (2025) -
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
por: Xie, Shuo, et al.
Publicado: (2025) -
Landscape-Aware Growing: The Power of a Little LAG
por: Karp, Stefani, et al.
Publicado: (2024) -
On the Role of Depth and Looping for In-Context Learning with Task Diversity
por: Gatmiry, Khashayar, et al.
Publicado: (2024) -
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
por: Gatmiry, Khashayar, et al.
Publicado: (2024)