PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Haonan, Chen, Brian, Li, Siquan, Liang, Xinhe, Lee, Hwee Kuan, Kawaguchi, Kenji, Hu, Tianyang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913043593560064
author Wang, Haonan
Chen, Brian
Li, Siquan
Liang, Xinhe
Lee, Hwee Kuan
Kawaguchi, Kenji
Hu, Tianyang
author_facet Wang, Haonan
Chen, Brian
Li, Siquan
Liang, Xinhe
Lee, Hwee Kuan
Kawaguchi, Kenji
Hu, Tianyang
contents Parameter-Efficient Fine-Tuning (PEFT) methods have become crucial for rapidly adapting large language models (LLMs) to downstream tasks. Prefix-Tuning, an early and effective PEFT technique, demonstrated the ability to achieve performance comparable to full fine-tuning with significantly reduced computational and memory overhead. However, despite its earlier success, its effectiveness in training modern state-of-the-art LLMs has been very limited. In this work, we demonstrate empirically that prefix-tuning underperforms on LLMs because of an inherent tradeoff between the contribution of the input prompt and the parameterized prefix within the attention head. This motivates us to introduce PrefixMemory-Tuning, an architecture that generalizes the principles of prefix-tuning while addressing its shortcomings by shifting the prefix module out of the attention head itself and improving its expressiveness. Our experiments show that, across diverse benchmarks, PrefixMemory-Tuning consistently outperforms existing prefix-tuning methods. Notably, it achieves competitive performance with modern PEFTs on several general benchmarks, highlighting a potential extension of prefix-tuning approaches to become state-of-the-art. Our findings suggest that by overcoming its inherent limitations, prefix-tuning can remain a competitive and relevant research direction in the landscape of parameter-efficient LLM adaptation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13674
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention
Wang, Haonan
Chen, Brian
Li, Siquan
Liang, Xinhe
Lee, Hwee Kuan
Kawaguchi, Kenji
Hu, Tianyang
Computation and Language
Artificial Intelligence
Parameter-Efficient Fine-Tuning (PEFT) methods have become crucial for rapidly adapting large language models (LLMs) to downstream tasks. Prefix-Tuning, an early and effective PEFT technique, demonstrated the ability to achieve performance comparable to full fine-tuning with significantly reduced computational and memory overhead. However, despite its earlier success, its effectiveness in training modern state-of-the-art LLMs has been very limited. In this work, we demonstrate empirically that prefix-tuning underperforms on LLMs because of an inherent tradeoff between the contribution of the input prompt and the parameterized prefix within the attention head. This motivates us to introduce PrefixMemory-Tuning, an architecture that generalizes the principles of prefix-tuning while addressing its shortcomings by shifting the prefix module out of the attention head itself and improving its expressiveness. Our experiments show that, across diverse benchmarks, PrefixMemory-Tuning consistently outperforms existing prefix-tuning methods. Notably, it achieves competitive performance with modern PEFTs on several general benchmarks, highlighting a potential extension of prefix-tuning approaches to become state-of-the-art. Our findings suggest that by overcoming its inherent limitations, prefix-tuning can remain a competitive and relevant research direction in the landscape of parameter-efficient LLM adaptation.
title PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.13674