DiaBlo: Diagonal Blocks Are Sufficient For Finetuning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gurses, Selcuk, Zhang, Aozhong, Deng, Yanxia, Dong, Xun, Li, Xin, Wang, Naigang, Yin, Penghang, Yang, Zi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912937834184704
author Gurses, Selcuk
Zhang, Aozhong
Deng, Yanxia
Dong, Xun
Li, Xin
Wang, Naigang
Yin, Penghang
Yang, Zi
author_facet Gurses, Selcuk
Zhang, Aozhong
Deng, Yanxia
Dong, Xun
Li, Xin
Wang, Naigang
Yin, Penghang
Yang, Zi
contents Fine-tuning is a critical step for adapting large language models (LLMs) to domain-specific downstream tasks. To mitigate the substantial computational and memory costs of full-model fine-tuning, Parameter-Efficient Fine-Tuning (PEFT) methods have been proposed to update only a small subset of model parameters. However, performance gaps between PEFT approaches and full-model fine-tuning still exist. In this work, we present DiaBlo, a simple yet effective PEFT approach that updates only the diagonal blocks of selected model weight matrices. Unlike Low-Rank Adaptation (LoRA) and its variants, DiaBlo eliminates the need for low-rank matrix products, thereby avoiding the reliance on auxiliary initialization schemes or customized optimization strategies to improve convergence. This design leads to stable and robust convergence while maintaining comparable memory efficiency and training speed to LoRA. Moreover, we provide theoretical guarantees showing that, under mild low-rank conditions, DiaBlo is more expressive than LoRA in the linear problem and converges to a stationary point of the general nonlinear full fine-tuning. Through extensive experiments across a range of tasks, including commonsense reasoning, arithmetic reasoning, code generation, and safety alignment, we show that fine-tuning only diagonal blocks is sufficient for strong and consistent performance. DiaBlo not only achieves competitive accuracy but also preserves high memory efficiency and fast fine-tuning speed. Codes are available at https://github.com/ziyangjoy/DiaBlo.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03230
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiaBlo: Diagonal Blocks Are Sufficient For Finetuning
Gurses, Selcuk
Zhang, Aozhong
Deng, Yanxia
Dong, Xun
Li, Xin
Wang, Naigang
Yin, Penghang
Yang, Zi
Machine Learning
Artificial Intelligence
Computation and Language
Optimization and Control
Fine-tuning is a critical step for adapting large language models (LLMs) to domain-specific downstream tasks. To mitigate the substantial computational and memory costs of full-model fine-tuning, Parameter-Efficient Fine-Tuning (PEFT) methods have been proposed to update only a small subset of model parameters. However, performance gaps between PEFT approaches and full-model fine-tuning still exist. In this work, we present DiaBlo, a simple yet effective PEFT approach that updates only the diagonal blocks of selected model weight matrices. Unlike Low-Rank Adaptation (LoRA) and its variants, DiaBlo eliminates the need for low-rank matrix products, thereby avoiding the reliance on auxiliary initialization schemes or customized optimization strategies to improve convergence. This design leads to stable and robust convergence while maintaining comparable memory efficiency and training speed to LoRA. Moreover, we provide theoretical guarantees showing that, under mild low-rank conditions, DiaBlo is more expressive than LoRA in the linear problem and converges to a stationary point of the general nonlinear full fine-tuning. Through extensive experiments across a range of tasks, including commonsense reasoning, arithmetic reasoning, code generation, and safety alignment, we show that fine-tuning only diagonal blocks is sufficient for strong and consistent performance. DiaBlo not only achieves competitive accuracy but also preserves high memory efficiency and fast fine-tuning speed. Codes are available at https://github.com/ziyangjoy/DiaBlo.
title DiaBlo: Diagonal Blocks Are Sufficient For Finetuning
topic Machine Learning
Artificial Intelligence
Computation and Language
Optimization and Control
url https://arxiv.org/abs/2506.03230