Towards Optimal Adapter Placement for Efficient Transfer Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nowak, Aleksandra I., Mercea, Otniel-Bogdan, Arnab, Anurag, Pfeiffer, Jonas, Dauphin, Yann, Evci, Utku
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910658164948992
author Nowak, Aleksandra I.
Mercea, Otniel-Bogdan
Arnab, Anurag
Pfeiffer, Jonas
Dauphin, Yann
Evci, Utku
author_facet Nowak, Aleksandra I.
Mercea, Otniel-Bogdan
Arnab, Anurag
Pfeiffer, Jonas
Dauphin, Yann
Evci, Utku
contents Parameter-efficient transfer learning (PETL) aims to adapt pre-trained models to new downstream tasks while minimizing the number of fine-tuned parameters. Adapters, a popular approach in PETL, inject additional capacity into existing networks by incorporating low-rank projections, achieving performance comparable to full fine-tuning with significantly fewer parameters. This paper investigates the relationship between the placement of an adapter and its performance. We observe that adapter location within a network significantly impacts its effectiveness, and that the optimal placement is task-dependent. To exploit this observation, we introduce an extended search space of adapter connections, including long-range and recurrent adapters. We demonstrate that even randomly selected adapter placements from this expanded space yield improved results, and that high-performing placements often correlate with high gradient rank. Our findings reveal that a small number of strategically placed adapters can match or exceed the performance of the common baseline of adding adapters in every block, opening a new avenue for research into optimal adapter placement strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15858
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Optimal Adapter Placement for Efficient Transfer Learning
Nowak, Aleksandra I.
Mercea, Otniel-Bogdan
Arnab, Anurag
Pfeiffer, Jonas
Dauphin, Yann
Evci, Utku
Machine Learning
Parameter-efficient transfer learning (PETL) aims to adapt pre-trained models to new downstream tasks while minimizing the number of fine-tuned parameters. Adapters, a popular approach in PETL, inject additional capacity into existing networks by incorporating low-rank projections, achieving performance comparable to full fine-tuning with significantly fewer parameters. This paper investigates the relationship between the placement of an adapter and its performance. We observe that adapter location within a network significantly impacts its effectiveness, and that the optimal placement is task-dependent. To exploit this observation, we introduce an extended search space of adapter connections, including long-range and recurrent adapters. We demonstrate that even randomly selected adapter placements from this expanded space yield improved results, and that high-performing placements often correlate with high gradient rank. Our findings reveal that a small number of strategically placed adapters can match or exceed the performance of the common baseline of adding adapters in every block, opening a new avenue for research into optimal adapter placement strategies.
title Towards Optimal Adapter Placement for Efficient Transfer Learning
topic Machine Learning
url https://arxiv.org/abs/2410.15858