Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Petrov, Egor, Evseev, Grigoriy, Antonov, Aleksey, Veprikov, Andrey, Bushkov, Nikolay, Moiseev, Stanislav, Beznosikov, Aleksandr
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917017628442624
author Petrov, Egor
Evseev, Grigoriy
Antonov, Aleksey
Veprikov, Andrey
Bushkov, Nikolay
Moiseev, Stanislav
Beznosikov, Aleksandr
author_facet Petrov, Egor
Evseev, Grigoriy
Antonov, Aleksey
Veprikov, Andrey
Bushkov, Nikolay
Moiseev, Stanislav
Beznosikov, Aleksandr
contents Fine-tuning Large Language Models (LLMs) is essential for adapting pre-trained models to downstream tasks. Yet traditional first-order optimizers such as Stochastic Gradient Descent (SGD) and Adam incur prohibitive memory and computational costs that scale poorly with model size. In this paper, we investigate zero-order (ZO) optimization methods as a memory- and compute-efficient alternative, particularly in the context of parameter-efficient fine-tuning techniques like LoRA. We propose $\texttt{JAGUAR SignSGD}$, a ZO momentum-based algorithm that extends ZO SignSGD, requiring the same number of parameters as the standard ZO SGD and only $\mathcal{O}(1)$ function evaluations per iteration. To the best of our knowledge, this is the first study to establish rigorous convergence guarantees for SignSGD in the stochastic ZO case. We further propose $\texttt{JAGUAR Muon}$, a novel ZO extension of the Muon optimizer that leverages the matrix structure of model parameters, and we provide its convergence rate under arbitrary stochastic noise. Through extensive experiments on challenging LLM fine-tuning benchmarks, we demonstrate that the proposed algorithms meet or exceed the convergence quality of standard first-order methods, achieving significant memory reduction. Our theoretical and empirical results establish new ZO optimization methods as a practical and theoretically grounded approach for resource-constrained LLM adaptation. Our code is available at https://github.com/brain-mmo-lab/ZO_LLM
format Preprint
id arxiv_https___arxiv_org_abs_2506_04430
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
Petrov, Egor
Evseev, Grigoriy
Antonov, Aleksey
Veprikov, Andrey
Bushkov, Nikolay
Moiseev, Stanislav
Beznosikov, Aleksandr
Machine Learning
Optimization and Control
Fine-tuning Large Language Models (LLMs) is essential for adapting pre-trained models to downstream tasks. Yet traditional first-order optimizers such as Stochastic Gradient Descent (SGD) and Adam incur prohibitive memory and computational costs that scale poorly with model size. In this paper, we investigate zero-order (ZO) optimization methods as a memory- and compute-efficient alternative, particularly in the context of parameter-efficient fine-tuning techniques like LoRA. We propose $\texttt{JAGUAR SignSGD}$, a ZO momentum-based algorithm that extends ZO SignSGD, requiring the same number of parameters as the standard ZO SGD and only $\mathcal{O}(1)$ function evaluations per iteration. To the best of our knowledge, this is the first study to establish rigorous convergence guarantees for SignSGD in the stochastic ZO case. We further propose $\texttt{JAGUAR Muon}$, a novel ZO extension of the Muon optimizer that leverages the matrix structure of model parameters, and we provide its convergence rate under arbitrary stochastic noise. Through extensive experiments on challenging LLM fine-tuning benchmarks, we demonstrate that the proposed algorithms meet or exceed the convergence quality of standard first-order methods, achieving significant memory reduction. Our theoretical and empirical results establish new ZO optimization methods as a practical and theoretically grounded approach for resource-constrained LLM adaptation. Our code is available at https://github.com/brain-mmo-lab/ZO_LLM
title Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2506.04430