MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Wei, Yaxiang, Zhang, Huang, Minhui, Xu, Mengfan, Zhang, Jiawei, Shen, Cong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917435147288576
author Shen, Wei
Yaxiang, Zhang
Huang, Minhui
Xu, Mengfan
Zhang, Jiawei
Shen, Cong
author_facet Shen, Wei
Yaxiang, Zhang
Huang, Minhui
Xu, Mengfan
Zhang, Jiawei
Shen, Cong
contents With increasing size of large language models (LLMs), full-parameter fine-tuning imposes substantial memory demands. To alleviate this, we propose a novel memory-efficient training paradigm called Momentum Low-rank compression (MLorc). The key idea of MLorc is to compress and reconstruct the momentum of matrix parameters during training to reduce memory consumption. Compared to LoRA, MLorc avoids enforcing a fixed-rank constraint on weight update matrices and thus enables full-parameter learning. Compared to GaLore, MLorc directly compress the momentum rather than gradients, thereby better preserving the training dynamics of full-parameter fine-tuning. We provide a theoretical guarantee for its convergence under mild assumptions. Empirically, MLorc consistently outperforms other memory-efficient training methods, matches or even exceeds the performance of full fine-tuning at small ranks (e.g., $r=4$), and generalizes well across different optimizers, all while not compromising time or memory efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01897
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
Shen, Wei
Yaxiang, Zhang
Huang, Minhui
Xu, Mengfan
Zhang, Jiawei
Shen, Cong
Machine Learning
Information Theory
Optimization and Control
With increasing size of large language models (LLMs), full-parameter fine-tuning imposes substantial memory demands. To alleviate this, we propose a novel memory-efficient training paradigm called Momentum Low-rank compression (MLorc). The key idea of MLorc is to compress and reconstruct the momentum of matrix parameters during training to reduce memory consumption. Compared to LoRA, MLorc avoids enforcing a fixed-rank constraint on weight update matrices and thus enables full-parameter learning. Compared to GaLore, MLorc directly compress the momentum rather than gradients, thereby better preserving the training dynamics of full-parameter fine-tuning. We provide a theoretical guarantee for its convergence under mild assumptions. Empirically, MLorc consistently outperforms other memory-efficient training methods, matches or even exceeds the performance of full fine-tuning at small ranks (e.g., $r=4$), and generalizes well across different optimizers, all while not compromising time or memory efficiency.
title MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
topic Machine Learning
Information Theory
Optimization and Control
url https://arxiv.org/abs/2506.01897