Linear Chain Transformation: Expanding Optimization Dynamics for Fine-Tuning Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yulong, Zuo, Chang, Xuan, Yin, Li, Hong, Wei, Ni
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915001466355712
author Wang, Yulong
Zuo, Chang
Xuan, Yin
Li, Hong
Wei, Ni
author_facet Wang, Yulong
Zuo, Chang
Xuan, Yin
Li, Hong
Wei, Ni
contents Fine-tuning large language models (LLMs) has become essential for adapting pretrained models to specific downstream tasks. In this paper, we propose Linear Chain Transformation (LinChain), a novel approach that introduces a sequence of linear transformations during fine-tuning to enrich optimization dynamics. By incorporating multiple linear transformations into the parameter update process, LinChain expands the effective rank of updates and enhances the model's ability to learn complex task-specific representations. We demonstrate that this method significantly improves the performance of LLM fine-tuning over state-of-the-art methods by providing more flexible optimization paths during training, while maintaining the inference efficiency of the resulting model. Our experiments on various benchmark tasks show that LinChain leads to better generalization, fewer learnable parameters, and improved task adaptation, making it a compelling strategy for LLM fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00039
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Linear Chain Transformation: Expanding Optimization Dynamics for Fine-Tuning Large Language Models
Wang, Yulong
Zuo, Chang
Xuan, Yin
Li, Hong
Wei, Ni
Computation and Language
Artificial Intelligence
Machine Learning
Fine-tuning large language models (LLMs) has become essential for adapting pretrained models to specific downstream tasks. In this paper, we propose Linear Chain Transformation (LinChain), a novel approach that introduces a sequence of linear transformations during fine-tuning to enrich optimization dynamics. By incorporating multiple linear transformations into the parameter update process, LinChain expands the effective rank of updates and enhances the model's ability to learn complex task-specific representations. We demonstrate that this method significantly improves the performance of LLM fine-tuning over state-of-the-art methods by providing more flexible optimization paths during training, while maintaining the inference efficiency of the resulting model. Our experiments on various benchmark tasks show that LinChain leads to better generalization, fewer learnable parameters, and improved task adaptation, making it a compelling strategy for LLM fine-tuning.
title Linear Chain Transformation: Expanding Optimization Dynamics for Fine-Tuning Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.00039