FZOO: Fast Zeroth-Order Optimizer for Fine-Tuning Large Language Models towards Adam-Scale Speed
Fuente:
arXiv
Saved in:
| Main Authors: | Dang, Sizhe, Guo, Yangyang, Zhao, Yanjun, Ye, Haishan, Zheng, Xiaodong, Dai, Guang, Tsang, Ivor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Second-Order Fine-Tuning without Pain for LLMs:A Hessian Informed Zeroth-Order Optimizer
by: Zhao, Yanjun, et al.
Published: (2024)
by: Zhao, Yanjun, et al.
Published: (2024)
ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning
by: Sun, Zhishen, et al.
Published: (2026)
by: Sun, Zhishen, et al.
Published: (2026)
From $O(mn)$ to $O(r^2)$: Two-Sided Low-Rank Communication for Adam in Distributed Training with Memory Efficiency
by: Dang, Sizhe, et al.
Published: (2026)
by: Dang, Sizhe, et al.
Published: (2026)
Numerical Sensitivity and Robustness: Exploring the Flaws of Mathematical Reasoning in Large Language Models
by: Sun, Zhishen, et al.
Published: (2025)
by: Sun, Zhishen, et al.
Published: (2025)
Double Variance Reduction: A Smoothing Trick for Composite Optimization Problems without First-Order Gradient
by: Di, Hao, et al.
Published: (2024)
by: Di, Hao, et al.
Published: (2024)
Why Does Adaptive Zeroth-Order Optimization Work?
by: Ye, Haishan, et al.
Published: (2026)
by: Ye, Haishan, et al.
Published: (2026)
Learning a Zeroth-Order Optimizer for Fine-Tuning LLMs
by: Zhang, Kairun, et al.
Published: (2025)
by: Zhang, Kairun, et al.
Published: (2025)
High-Probability Guarantees for Random Zeroth-Order (Stochastic) Gradient Descent
by: Ye, Haishan
Published: (2026)
by: Ye, Haishan
Published: (2026)
High-Probability Guarantees for Random Zeroth-Order Gradient Descent on Smooth Functions
by: Ye, Haishan
Published: (2026)
by: Ye, Haishan
Published: (2026)
A Unified Zeroth-Order Optimization Framework via Oblivious Randomized Sketching
by: Ye, Haishan, et al.
Published: (2025)
by: Ye, Haishan, et al.
Published: (2025)
Zeroth-Order Fine-Tuning of LLMs with Extreme Sparsity
by: Guo, Wentao, et al.
Published: (2024)
by: Guo, Wentao, et al.
Published: (2024)
An Enhanced Zeroth-Order Stochastic Frank-Wolfe Framework for Constrained Finite-Sum Optimization
by: Ye, Haishan, et al.
Published: (2025)
by: Ye, Haishan, et al.
Published: (2025)
Simultaneous Computation and Memory Efficient Zeroth-Order Optimizer for Fine-Tuning Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
OAT-Rephrase: Optimization-Aware Training Data Rephrasing for Zeroth-Order LLM Fine-Tuning
by: Long, Jikai, et al.
Published: (2025)
by: Long, Jikai, et al.
Published: (2025)
Explicit and Non-asymptotic Query Complexities of Rank-Based Zeroth-order Algorithms on Smooth Functions
by: Ye, Haishan
Published: (2025)
by: Ye, Haishan
Published: (2025)
Riemannian Momentum Tracking: Distributed Optimization with Momentum on Compact Submanifolds
by: Chen, Jun, et al.
Published: (2026)
by: Chen, Jun, et al.
Published: (2026)
Privacy Leaks by Adversaries: Adversarial Iterations for Membership Inference Attack
by: Xue, Jing, et al.
Published: (2025)
by: Xue, Jing, et al.
Published: (2025)
Explicit and Non-asymptotic Query Complexities of Rank-Based Zeroth-order Algorithm on Stochastic Smooth Functions
by: Ye, Haishan
Published: (2025)
by: Ye, Haishan
Published: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
MaZO: Masked Zeroth-Order Optimization for Multi-Task Fine-Tuning of Large Language Models
by: Zhang, Zhen, et al.
Published: (2025)
by: Zhang, Zhen, et al.
Published: (2025)
QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language Models
by: Zhou, Jiajun, et al.
Published: (2025)
by: Zhou, Jiajun, et al.
Published: (2025)
On-Device Fine-Tuning via Backprop-Free Zeroth-Order Optimization
by: Katti, Prabodh, et al.
Published: (2025)
by: Katti, Prabodh, et al.
Published: (2025)
Low-Rank Curvature for Zeroth-Order Optimization in LLM Fine-Tuning
by: Seung, Hyunseok, et al.
Published: (2025)
by: Seung, Hyunseok, et al.
Published: (2025)
Decentralized Riemannian Conjugate Gradient Method on the Stiefel Manifold
by: Chen, Jun, et al.
Published: (2023)
by: Chen, Jun, et al.
Published: (2023)
Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A Benchmark
by: Zhang, Yihua, et al.
Published: (2024)
by: Zhang, Yihua, et al.
Published: (2024)
Mitigating Non-IID Drift in Zeroth-Order Federated LLM Fine-Tuning with Transferable Sparsity
by: Ran, Yide, et al.
Published: (2025)
by: Ran, Yide, et al.
Published: (2025)
LLM Zeroth-Order Fine-Tuning is an Inference Workload
by: Li, Zelin, et al.
Published: (2026)
by: Li, Zelin, et al.
Published: (2026)
Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
by: Yu, Ziming, et al.
Published: (2024)
by: Yu, Ziming, et al.
Published: (2024)
AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments
by: Cai, Zhijie, et al.
Published: (2026)
by: Cai, Zhijie, et al.
Published: (2026)
Can a One-Point Feedback Zeroth-order Algorithm Achieve Linear Dimension Dependent Sample Complexity?
by: Ye, Haishan, et al.
Published: (2025)
by: Ye, Haishan, et al.
Published: (2025)
On the Convergence of Single-Loop Stochastic Bilevel Optimization with Approximate Implicit Differentiation
by: Zhou, Yubo, et al.
Published: (2026)
by: Zhou, Yubo, et al.
Published: (2026)
Less is more: Embracing sparsity and interpolation with Esiformer for time series forecasting
by: Guo, Yangyang, et al.
Published: (2024)
by: Guo, Yangyang, et al.
Published: (2024)
Robust and Efficient Zeroth-Order LLM Fine-Tuning via Adaptive Bayesian Subspace Optimizer
by: Feng, Jian, et al.
Published: (2026)
by: Feng, Jian, et al.
Published: (2026)
A First-Order Multi-Gradient Algorithm for Multi-Objective Bi-Level Optimization
by: Ye, Feiyang, et al.
Published: (2024)
by: Ye, Feiyang, et al.
Published: (2024)
RealDiffusion: Physics-informed Attention for Multi-character Storybook Generation
by: Zhao, Qi, et al.
Published: (2026)
by: Zhao, Qi, et al.
Published: (2026)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
by: Gautam, Tanmay, et al.
Published: (2024)
by: Gautam, Tanmay, et al.
Published: (2024)
MSCR: Exploring the Vulnerability of LLMs' Mathematical Reasoning Abilities Using Multi-Source Candidate Replacement
by: Sun, Zhishen, et al.
Published: (2025)
by: Sun, Zhishen, et al.
Published: (2025)
AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-Tuning
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Refining Adaptive Zeroth-Order Optimization at Ease
by: Shu, Yao, et al.
Published: (2025)
by: Shu, Yao, et al.
Published: (2025)
Prior-Informed Zeroth-Order Optimization with Adaptive Direction Alignment for Memory-Efficient LLM Fine-Tuning
by: Jin, Feihu, et al.
Published: (2026)
by: Jin, Feihu, et al.
Published: (2026)
Similar Items
-
Second-Order Fine-Tuning without Pain for LLMs:A Hessian Informed Zeroth-Order Optimizer
by: Zhao, Yanjun, et al.
Published: (2024) -
ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning
by: Sun, Zhishen, et al.
Published: (2026) -
From $O(mn)$ to $O(r^2)$: Two-Sided Low-Rank Communication for Adam in Distributed Training with Memory Efficiency
by: Dang, Sizhe, et al.
Published: (2026) -
Numerical Sensitivity and Robustness: Exploring the Flaws of Mathematical Reasoning in Large Language Models
by: Sun, Zhishen, et al.
Published: (2025) -
Double Variance Reduction: A Smoothing Trick for Composite Optimization Problems without First-Order Gradient
by: Di, Hao, et al.
Published: (2024)