Riemannian Preconditioned LoRA for Fine-Tuning Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Fangzhao, Pilanci, Mert |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyzing Neural Network-Based Generative Diffusion Models through Convex Optimization
by: Zhang, Fangzhao, et al.
Published: (2024)
by: Zhang, Fangzhao, et al.
Published: (2024)
Active Learning of Deep Neural Networks via Gradient-Free Cutting Planes
by: Zhang, Erica, et al.
Published: (2024)
by: Zhang, Erica, et al.
Published: (2024)
Towards Quantifying the Preconditioning Effect of Adam
by: Das, Rudrajit, et al.
Published: (2024)
by: Das, Rudrajit, et al.
Published: (2024)
Newton Meets Marchenko-Pastur: Massively Parallel Second-Order Optimization with Hessian Sketching and Debiasing
by: Romanov, Elad, et al.
Published: (2024)
by: Romanov, Elad, et al.
Published: (2024)
Fast and Provable Tensor-Train Format Tensor Completion via Precondtioned Riemannian Gradient Descent
by: Bian, Fengmiao, et al.
Published: (2025)
by: Bian, Fengmiao, et al.
Published: (2025)
Natural Riemannian gradient for learning functional tensor networks
by: Klug, Nikolas, et al.
Published: (2026)
by: Klug, Nikolas, et al.
Published: (2026)
Why is Normalization Preferred? A Worst-Case Complexity Theory for Stochastically Preconditioned SGD under Heavy-Tailed Noise
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
Optimal Shrinkage for Distributed Second-Order Optimization
by: Zhang, Fangzhao, et al.
Published: (2024)
by: Zhang, Fangzhao, et al.
Published: (2024)
Machine Learning and Control: Foundations, Advances, and Perspectives
by: Zuazua, Enrique
Published: (2025)
by: Zuazua, Enrique
Published: (2025)
Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression
by: Yan, Mingsong, et al.
Published: (2026)
by: Yan, Mingsong, et al.
Published: (2026)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
Solving Dense Linear Systems Faster Than via Preconditioning
by: Dereziński, Michał, et al.
Published: (2023)
by: Dereziński, Michał, et al.
Published: (2023)
WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle Points
by: Li, Dongyue, et al.
Published: (2026)
by: Li, Dongyue, et al.
Published: (2026)
Modified Loss of Momentum Gradient Descent: Fine-Grained Analysis
by: Cattaneo, Matias D., et al.
Published: (2025)
by: Cattaneo, Matias D., et al.
Published: (2025)
Faster Linear Systems and Matrix Norm Approximation via Multi-level Sketched Preconditioning
by: Dereziński, Michał, et al.
Published: (2024)
by: Dereziński, Michał, et al.
Published: (2024)
Subhomogeneous Deep Equilibrium Models
by: Sittoni, Pietro, et al.
Published: (2024)
by: Sittoni, Pietro, et al.
Published: (2024)
Multi-level Optimal Control with Neural Surrogate Models
by: Kalise, Dante, et al.
Published: (2024)
by: Kalise, Dante, et al.
Published: (2024)
Modeling AdaGrad, RMSProp, and Adam with Integro-Differential Equations
by: Heredia, Carlos
Published: (2024)
by: Heredia, Carlos
Published: (2024)
Efficient Algorithms for Regularized Nonnegative Scale-invariant Low-rank Approximation Models
by: Cohen, Jeremy E., et al.
Published: (2024)
by: Cohen, Jeremy E., et al.
Published: (2024)
Spectral Adapter: Fine-Tuning in Spectral Space
by: Zhang, Fangzhao, et al.
Published: (2024)
by: Zhang, Fangzhao, et al.
Published: (2024)
Scalable Approximate Optimal Diagonal Preconditioning
by: Gao, Wenzhi, et al.
Published: (2023)
by: Gao, Wenzhi, et al.
Published: (2023)
Accuracy of Discretely Sampled Stochastic Policies in Continuous-time Reinforcement Learning
by: Jia, Yanwei, et al.
Published: (2025)
by: Jia, Yanwei, et al.
Published: (2025)
Enhanced Adaptive Gradient Algorithms for Nonconvex-PL Minimax Optimization
by: Huang, Feihu, et al.
Published: (2023)
by: Huang, Feihu, et al.
Published: (2023)
Nonlinear Assimilation via Score-based Sequential Langevin Sampling
by: Ding, Zhao, et al.
Published: (2024)
by: Ding, Zhao, et al.
Published: (2024)
Preconditioning transformations of adjoint systems for evolution equations
by: Tran, Brian K., et al.
Published: (2025)
by: Tran, Brian K., et al.
Published: (2025)
Convex Relaxations of ReLU Neural Networks Approximate Global Optima in Polynomial Time
by: Kim, Sungyoon, et al.
Published: (2024)
by: Kim, Sungyoon, et al.
Published: (2024)
Learning Regularization Functionals for Inverse Problems: A Comparative Study
by: Hertrich, Johannes, et al.
Published: (2025)
by: Hertrich, Johannes, et al.
Published: (2025)
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
On the numerical reliability of nonsmooth autodiff: a MaxPool case study
by: Boustany, Ryan
Published: (2024)
by: Boustany, Ryan
Published: (2024)
A Gauss-Newton Approach for Min-Max Optimization in Generative Adversarial Networks
by: Mishra, Neel, et al.
Published: (2024)
by: Mishra, Neel, et al.
Published: (2024)
Flattened one-bit stochastic gradient descent: compressed distributed optimization with controlled variance
by: Stollenwerk, Alexander, et al.
Published: (2024)
by: Stollenwerk, Alexander, et al.
Published: (2024)
Efficient Trajectory Inference in Wasserstein Space Using Consecutive Averaging
by: Banerjee, Amartya, et al.
Published: (2024)
by: Banerjee, Amartya, et al.
Published: (2024)
Anderson Acceleration in Nonsmooth Problems: Local Convergence via Active Manifold Identification
by: Li, Kexin, et al.
Published: (2024)
by: Li, Kexin, et al.
Published: (2024)
KANtrol: A Physics-Informed Kolmogorov-Arnold Network Framework for Solving Multi-Dimensional and Fractional Optimal Control Problems
by: Aghaei, Alireza Afzal
Published: (2024)
by: Aghaei, Alireza Afzal
Published: (2024)
Real-time optimal control of high-dimensional parametrized systems by deep learning-based reduced order models
by: Tomasetto, Matteo, et al.
Published: (2024)
by: Tomasetto, Matteo, et al.
Published: (2024)
Cubic regularized subspace Newton for non-convex optimization
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Super Gradient Descent: Global Optimization requires Global Gradient
by: Achour, Seifeddine
Published: (2024)
by: Achour, Seifeddine
Published: (2024)
Latent feedback control of distributed systems in multiple scenarios through deep learning-based reduced order models
by: Tomasetto, Matteo, et al.
Published: (2024)
by: Tomasetto, Matteo, et al.
Published: (2024)
Quantitative Convergences of Lie Group Momentum Optimizers
by: Kong, Lingkai, et al.
Published: (2024)
by: Kong, Lingkai, et al.
Published: (2024)
Learning incomplete factorization preconditioners for GMRES
by: Häusner, Paul, et al.
Published: (2024)
by: Häusner, Paul, et al.
Published: (2024)
Similar Items
-
Analyzing Neural Network-Based Generative Diffusion Models through Convex Optimization
by: Zhang, Fangzhao, et al.
Published: (2024) -
Active Learning of Deep Neural Networks via Gradient-Free Cutting Planes
by: Zhang, Erica, et al.
Published: (2024) -
Towards Quantifying the Preconditioning Effect of Adam
by: Das, Rudrajit, et al.
Published: (2024) -
Newton Meets Marchenko-Pastur: Massively Parallel Second-Order Optimization with Hessian Sketching and Debiasing
by: Romanov, Elad, et al.
Published: (2024) -
Fast and Provable Tensor-Train Format Tensor Completion via Precondtioned Riemannian Gradient Descent
by: Bian, Fengmiao, et al.
Published: (2025)