DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hu, Xiaolin, Cheng, Xiang, Liu, Peiyu, Liu, Wei, Luan, Jian, Wang, Bin, Liu, Yong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915085042057216
author Hu, Xiaolin
Cheng, Xiang
Liu, Peiyu
Liu, Wei
Luan, Jian
Wang, Bin
Liu, Yong
author_facet Hu, Xiaolin
Cheng, Xiang
Liu, Peiyu
Liu, Wei
Luan, Jian
Wang, Bin
Liu, Yong
contents Low-rank adaptation (LoRA) reduces the computational and memory demands of fine-tuning large language models (LLMs) by approximating updates with low-rank matrices. However, low-rank approximation in two-dimensional space fails to capture high-dimensional structures within the target matrix. Recently, tensor decomposition methods have been explored for fine-tuning LLMs, leveraging their ability to extract structured information. Yet, these approaches primarily rely on random initialization, and the impact of initialization on tensor adaptation remains underexplored. In this paper, we reveal that random initialization significantly diverges from the validation loss achieved by full fine-tuning. To address this, we propose Weight-Decomposed Tensor Adaptation (DoTA), which leverages the Matrix Product Operator (MPO) decomposition of pre-trained weights for effective initialization in fine-tuning LLMs. Additionally, we introduce QDoTA, a quantized version of DoTA designed for 4-bit quantization. Experiments on commonsense and arithmetic reasoning tasks show that DoTA outperforms random initialization methods with fewer parameters. QDoTA further reduces memory consumption and achieves comparable performance to DoTA on commonsense reasoning tasks. We will release our code to support future research.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20891
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models
Hu, Xiaolin
Cheng, Xiang
Liu, Peiyu
Liu, Wei
Luan, Jian
Wang, Bin
Liu, Yong
Computation and Language
Machine Learning
Low-rank adaptation (LoRA) reduces the computational and memory demands of fine-tuning large language models (LLMs) by approximating updates with low-rank matrices. However, low-rank approximation in two-dimensional space fails to capture high-dimensional structures within the target matrix. Recently, tensor decomposition methods have been explored for fine-tuning LLMs, leveraging their ability to extract structured information. Yet, these approaches primarily rely on random initialization, and the impact of initialization on tensor adaptation remains underexplored. In this paper, we reveal that random initialization significantly diverges from the validation loss achieved by full fine-tuning. To address this, we propose Weight-Decomposed Tensor Adaptation (DoTA), which leverages the Matrix Product Operator (MPO) decomposition of pre-trained weights for effective initialization in fine-tuning LLMs. Additionally, we introduce QDoTA, a quantized version of DoTA designed for 4-bit quantization. Experiments on commonsense and arithmetic reasoning tasks show that DoTA outperforms random initialization methods with fewer parameters. QDoTA further reduces memory consumption and achieves comparable performance to DoTA on commonsense reasoning tasks. We will release our code to support future research.
title DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.20891