Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuang, Zhan, Wang, Xiequn, Li, Wei, Zhang, Yulong, Huang, Qiushi, Chen, Shuhao, Wang, Xuehao, Wei, Yanbin, Nie, Yuhe, Ma, Kede, Zhang, Yu, Wei, Ying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909707991515136
author Zhuang, Zhan
Wang, Xiequn
Li, Wei
Zhang, Yulong
Huang, Qiushi
Chen, Shuhao
Wang, Xuehao
Wei, Yanbin
Nie, Yuhe
Ma, Kede
Zhang, Yu
Wei, Ying
author_facet Zhuang, Zhan
Wang, Xiequn
Li, Wei
Zhang, Yulong
Huang, Qiushi
Chen, Shuhao
Wang, Xuehao
Wei, Yanbin
Nie, Yuhe
Ma, Kede
Zhang, Yu
Wei, Ying
contents Low-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal minima near their initialization. This hampers model generalization and limits downstream operators such as adapter merging and pruning. Here, we propose CoTo, a progressive training strategy that gradually increases adapters' activation probability over the course of fine-tuning. By stochastically deactivating adapters, CoTo encourages more balanced optimization and broader exploration of the loss landscape. We provide a theoretical analysis showing that CoTo promotes layer-wise dropout stability and linear mode connectivity, and we adopt a cooperative-game approach to quantify each adapter's marginal contribution. Extensive experiments demonstrate that CoTo consistently boosts single-task performance, enhances multi-task merging accuracy, improves pruning robustness, and reduces training overhead, all while remaining compatible with diverse LoRA variants. Code is available at https://github.com/zwebzone/coto.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation
Zhuang, Zhan
Wang, Xiequn
Li, Wei
Zhang, Yulong
Huang, Qiushi
Chen, Shuhao
Wang, Xuehao
Wei, Yanbin
Nie, Yuhe
Ma, Kede
Zhang, Yu
Wei, Ying
Machine Learning
Low-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal minima near their initialization. This hampers model generalization and limits downstream operators such as adapter merging and pruning. Here, we propose CoTo, a progressive training strategy that gradually increases adapters' activation probability over the course of fine-tuning. By stochastically deactivating adapters, CoTo encourages more balanced optimization and broader exploration of the loss landscape. We provide a theoretical analysis showing that CoTo promotes layer-wise dropout stability and linear mode connectivity, and we adopt a cooperative-game approach to quantify each adapter's marginal contribution. Extensive experiments demonstrate that CoTo consistently boosts single-task performance, enhances multi-task merging accuracy, improves pruning robustness, and reduces training overhead, all while remaining compatible with diverse LoRA variants. Code is available at https://github.com/zwebzone/coto.
title Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation
topic Machine Learning
url https://arxiv.org/abs/2506.05713