Towards Optimal Adapter Placement for Efficient Transfer Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910658164948992 |
|---|---|
| author | Nowak, Aleksandra I. Mercea, Otniel-Bogdan Arnab, Anurag Pfeiffer, Jonas Dauphin, Yann Evci, Utku |
| author_facet | Nowak, Aleksandra I. Mercea, Otniel-Bogdan Arnab, Anurag Pfeiffer, Jonas Dauphin, Yann Evci, Utku |
| contents | Parameter-efficient transfer learning (PETL) aims to adapt pre-trained models to new downstream tasks while minimizing the number of fine-tuned parameters. Adapters, a popular approach in PETL, inject additional capacity into existing networks by incorporating low-rank projections, achieving performance comparable to full fine-tuning with significantly fewer parameters. This paper investigates the relationship between the placement of an adapter and its performance. We observe that adapter location within a network significantly impacts its effectiveness, and that the optimal placement is task-dependent. To exploit this observation, we introduce an extended search space of adapter connections, including long-range and recurrent adapters. We demonstrate that even randomly selected adapter placements from this expanded space yield improved results, and that high-performing placements often correlate with high gradient rank. Our findings reveal that a small number of strategically placed adapters can match or exceed the performance of the common baseline of adding adapters in every block, opening a new avenue for research into optimal adapter placement strategies. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_15858 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Towards Optimal Adapter Placement for Efficient Transfer Learning Nowak, Aleksandra I. Mercea, Otniel-Bogdan Arnab, Anurag Pfeiffer, Jonas Dauphin, Yann Evci, Utku Machine Learning Parameter-efficient transfer learning (PETL) aims to adapt pre-trained models to new downstream tasks while minimizing the number of fine-tuned parameters. Adapters, a popular approach in PETL, inject additional capacity into existing networks by incorporating low-rank projections, achieving performance comparable to full fine-tuning with significantly fewer parameters. This paper investigates the relationship between the placement of an adapter and its performance. We observe that adapter location within a network significantly impacts its effectiveness, and that the optimal placement is task-dependent. To exploit this observation, we introduce an extended search space of adapter connections, including long-range and recurrent adapters. We demonstrate that even randomly selected adapter placements from this expanded space yield improved results, and that high-performing placements often correlate with high gradient rank. Our findings reveal that a small number of strategically placed adapters can match or exceed the performance of the common baseline of adding adapters in every block, opening a new avenue for research into optimal adapter placement strategies. |
| title | Towards Optimal Adapter Placement for Efficient Transfer Learning |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2410.15858 |