AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiacheng, Tan, Jianchao, Yang, Zhidong, Huo, Feiye, Sun, Yerui, Xie, Yuchen, Cai, Xunliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914232651481088
author Li, Jiacheng
Tan, Jianchao
Yang, Zhidong
Huo, Feiye
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
author_facet Li, Jiacheng
Tan, Jianchao
Yang, Zhidong
Huo, Feiye
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
contents Low-Rank Adaptation (LoRA) is a widely adopted parameter-efficient fine-tuning (PEFT) method. However, its linear adaptation process limits its expressive power. This means there is a gap between the expressive power of linear training and non-linear training. To bridge this gap, we propose AFA-LoRA, a novel training strategy that brings non-linear expressivity to LoRA while maintaining its seamless mergeability. Our key innovation is an annealed activation function that transitions from a non-linear to a linear transformation during training, allowing the adapter to initially adopt stronger representational capabilities before converging to a mergeable linear form. We implement our method on supervised fine-tuning, reinforcement learning, and speculative decoding. The results show that AFA-LoRA reduces the performance gap between LoRA and full-parameter training. This work enables a more powerful and practical paradigm of parameter-efficient adaptation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22455
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing
Li, Jiacheng
Tan, Jianchao
Yang, Zhidong
Huo, Feiye
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
Machine Learning
Computation and Language
Low-Rank Adaptation (LoRA) is a widely adopted parameter-efficient fine-tuning (PEFT) method. However, its linear adaptation process limits its expressive power. This means there is a gap between the expressive power of linear training and non-linear training. To bridge this gap, we propose AFA-LoRA, a novel training strategy that brings non-linear expressivity to LoRA while maintaining its seamless mergeability. Our key innovation is an annealed activation function that transitions from a non-linear to a linear transformation during training, allowing the adapter to initially adopt stronger representational capabilities before converging to a mergeable linear form. We implement our method on supervised fine-tuning, reinforcement learning, and speculative decoding. The results show that AFA-LoRA reduces the performance gap between LoRA and full-parameter training. This work enables a more powerful and practical paradigm of parameter-efficient adaptation.
title AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2512.22455