LORENZA: Enhancing Generalization in Low-Rank Gradient LLM Training via Efficient Zeroth-Order Adaptive SAM
Fuente:
arXiv
Guardado en:
| Autores principales: | Refael, Yehonathan, Arbel, Iftach, Lindenbaum, Ofir, Tirer, Tom |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM Training
por: Refael, Yehonathan, et al.
Publicado: (2025)
por: Refael, Yehonathan, et al.
Publicado: (2025)
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
por: Arbel, Iftach, et al.
Publicado: (2024)
por: Arbel, Iftach, et al.
Publicado: (2024)
Unveiling Multiple Descents in Unsupervised Autoencoders
por: Rahimi, Kobi, et al.
Publicado: (2024)
por: Rahimi, Kobi, et al.
Publicado: (2024)
AdaRankGrad: Adaptive Gradient-Rank and Moments for Memory-Efficient LLMs Training and Fine-Tuning
por: Refael, Yehonathan, et al.
Publicado: (2024)
por: Refael, Yehonathan, et al.
Publicado: (2024)
Train Less, Infer Faster: Efficient Model Finetuning and Compression via Structured Sparsity
por: Svirsky, Jonathan, et al.
Publicado: (2026)
por: Svirsky, Jonathan, et al.
Publicado: (2026)
Learning k-Level Structured Sparse Neural Networks Using Group Envelope Regularization
por: Refael, Yehonathan, et al.
Publicado: (2022)
por: Refael, Yehonathan, et al.
Publicado: (2022)
On Adaptivity in Zeroth-Order Optimization
por: Dbouk, Hassan, et al.
Publicado: (2026)
por: Dbouk, Hassan, et al.
Publicado: (2026)
FineGates: LLMs Finetuning with Compression using Stochastic Gates
por: Svirsky, Jonathan, et al.
Publicado: (2024)
por: Svirsky, Jonathan, et al.
Publicado: (2024)
On the Inherent Privacy of Zeroth Order Projected Gradient Descent
por: Gupta, Devansh, et al.
Publicado: (2025)
por: Gupta, Devansh, et al.
Publicado: (2025)
Zeroth-Order Hard-Thresholding: Gradient Error vs. Expansivity
por: de Vazelhes, William, et al.
Publicado: (2022)
por: de Vazelhes, William, et al.
Publicado: (2022)
Riemannian Zeroth-Order Gradient Estimation with Structure-Preserving Metrics for Geodesically Incomplete Manifolds
por: Ma, Shaocong, et al.
Publicado: (2026)
por: Ma, Shaocong, et al.
Publicado: (2026)
Obtaining Lower Query Complexities through Lightweight Zeroth-Order Proximal Gradient Algorithms
por: Gu, Bin, et al.
Publicado: (2024)
por: Gu, Bin, et al.
Publicado: (2024)
Fully Adaptive Zeroth-Order Method for Minimizing Functions with Compressible Gradients
por: Grapiglia, Geovani Nunes, et al.
Publicado: (2025)
por: Grapiglia, Geovani Nunes, et al.
Publicado: (2025)
On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization
por: Ma, Shaocong, et al.
Publicado: (2025)
por: Ma, Shaocong, et al.
Publicado: (2025)
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
por: Chen, Jiahe, et al.
Publicado: (2025)
por: Chen, Jiahe, et al.
Publicado: (2025)
Fully Zeroth-Order Bilevel Programming via Gaussian Smoothing
por: Aghasi, Alireza, et al.
Publicado: (2024)
por: Aghasi, Alireza, et al.
Publicado: (2024)
Efficient Low-Tubal-Rank Tensor Estimation via Alternating Preconditioned Gradient Descent
por: Liu, Zhiyu, et al.
Publicado: (2025)
por: Liu, Zhiyu, et al.
Publicado: (2025)
Zeroth-Order primal-dual Alternating Projection Gradient Algorithms for Nonconvex Minimax Problems with Coupled linear Constraints
por: Zhang, Huiling, et al.
Publicado: (2024)
por: Zhang, Huiling, et al.
Publicado: (2024)
Low-Tubal-Rank Tensor Recovery via Factorized Gradient Descent
por: Liu, Zhiyu, et al.
Publicado: (2024)
por: Liu, Zhiyu, et al.
Publicado: (2024)
Zeroth-Order Methods for Stochastic Nonconvex Nonsmooth Composite Optimization
por: Chen, Ziyi, et al.
Publicado: (2025)
por: Chen, Ziyi, et al.
Publicado: (2025)
Zeroth-Order Optimization at the Edge of Stability
por: Song, Minhak, et al.
Publicado: (2026)
por: Song, Minhak, et al.
Publicado: (2026)
Why Does Adaptive Zeroth-Order Optimization Work?
por: Ye, Haishan, et al.
Publicado: (2026)
por: Ye, Haishan, et al.
Publicado: (2026)
A Randomized Zeroth-Order Hierarchical Framework for Heterogeneous Federated Learning
por: Qiu, Yuyang, et al.
Publicado: (2025)
por: Qiu, Yuyang, et al.
Publicado: (2025)
Minimisation of Polyak-Łojasewicz Functions Using Random Zeroth-Order Oracles
por: Farzin, Amir Ali, et al.
Publicado: (2024)
por: Farzin, Amir Ali, et al.
Publicado: (2024)
High-Probability Guarantees for Random Zeroth-Order (Stochastic) Gradient Descent
por: Ye, Haishan
Publicado: (2026)
por: Ye, Haishan
Publicado: (2026)
Distributed Zeroth-Order Optimization with Rademacher Perturbations and Momentum Gradient Tracking
por: Su, Yanxu, et al.
Publicado: (2026)
por: Su, Yanxu, et al.
Publicado: (2026)
Zeroth-Order Optimization Finds Flat Minima
por: Zhang, Liang, et al.
Publicado: (2025)
por: Zhang, Liang, et al.
Publicado: (2025)
Certified Multi-Fidelity Zeroth-Order Optimization
por: de Montbrun, Étienne, et al.
Publicado: (2023)
por: de Montbrun, Étienne, et al.
Publicado: (2023)
Private Zeroth-Order Nonsmooth Nonconvex Optimization
por: Zhang, Qinzi, et al.
Publicado: (2024)
por: Zhang, Qinzi, et al.
Publicado: (2024)
Zeroth-Order Stochastic Mirror Descent Algorithms for Minimax Excess Risk Optimization
por: Gu, Zhihao, et al.
Publicado: (2024)
por: Gu, Zhihao, et al.
Publicado: (2024)
Unbiased Gradient Low-Rank Projection
por: Pan, Rui, et al.
Publicado: (2025)
por: Pan, Rui, et al.
Publicado: (2025)
Greedy Low-Rank Gradient Compression for Distributed Learning with Convergence Guarantees
por: Chen, Chuyan, et al.
Publicado: (2025)
por: Chen, Chuyan, et al.
Publicado: (2025)
High-Probability Guarantees for Random Zeroth-Order Gradient Descent on Smooth Functions
por: Ye, Haishan
Publicado: (2026)
por: Ye, Haishan
Publicado: (2026)
A Zeroth-Order Extra-Gradient Method for Black-Box Constrained Optimization
por: Zhou, Yuke, et al.
Publicado: (2025)
por: Zhou, Yuke, et al.
Publicado: (2025)
Accelerating Single-Point Zeroth-Order Optimization with Regression-Based Gradient Surrogates
por: Chen, Xin, et al.
Publicado: (2025)
por: Chen, Xin, et al.
Publicado: (2025)
Asynchronous Distributed Reinforcement Learning for LQR Control via Zeroth-Order Block Coordinate Descent
por: Jing, Gangshan, et al.
Publicado: (2021)
por: Jing, Gangshan, et al.
Publicado: (2021)
Explicit and Non-asymptotic Query Complexities of Rank-Based Zeroth-order Algorithm on Stochastic Smooth Functions
por: Ye, Haishan
Publicado: (2025)
por: Ye, Haishan
Publicado: (2025)
Variance-Reduced Gradient Estimator for Nonconvex Zeroth-Order Distributed Optimization
por: Mu, Huaiyi, et al.
Publicado: (2024)
por: Mu, Huaiyi, et al.
Publicado: (2024)
Efficient Model-Free Exploration in Low-Rank MDPs
por: Mhammedi, Zakaria, et al.
Publicado: (2023)
por: Mhammedi, Zakaria, et al.
Publicado: (2023)
Fast and Accurate Estimation of Low-Rank Matrices from Noisy Measurements via Preconditioned Non-Convex Gradient Descent
por: Zhang, Gavin, et al.
Publicado: (2023)
por: Zhang, Gavin, et al.
Publicado: (2023)
Ejemplares similares
-
SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM Training
por: Refael, Yehonathan, et al.
Publicado: (2025) -
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
por: Arbel, Iftach, et al.
Publicado: (2024) -
Unveiling Multiple Descents in Unsupervised Autoencoders
por: Rahimi, Kobi, et al.
Publicado: (2024) -
AdaRankGrad: Adaptive Gradient-Rank and Moments for Memory-Efficient LLMs Training and Fine-Tuning
por: Refael, Yehonathan, et al.
Publicado: (2024) -
Train Less, Infer Faster: Efficient Model Finetuning and Compression via Structured Sparsity
por: Svirsky, Jonathan, et al.
Publicado: (2026)