PILLOW: Enhancing Efficient Instruction Fine-tuning via Prompt Matching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866909337891373056 |
|---|---|
| author | Qi, Zhenting Tan, Xiaoyu Shi, Shaojie Qu, Chao Xu, Yinghui Qi, Yuan |
| author_facet | Qi, Zhenting Tan, Xiaoyu Shi, Shaojie Qu, Chao Xu, Yinghui Qi, Yuan |
| contents | Instruction fine-tuning has conventionally been employed to adapt Large Language Models (LLMs) to a variety of tasks. Nonetheless, this technique often necessitates substantial computational resources, making it impractical for deployment by individuals or small-scale entities. Recently, Low-Rank Adaptation (LoRA) has become a promising alternative, offering high capabilities on par with full tuning with reduced resource overhead. However, attaining satisfactory performance through the fine-tuning of LoRA is a non-trivial challenge. In this paper, we propose PILLOW, which aims to improve LoRA's performance by a discrimination-based prompting method, leveraging LLMs' In-Context Learning ability. PILLOW incorporates a matching network that selects prompts from a user-defined prompt pool, concatenates the selected prompts with the user instruction as input, and performs inference using the LoRA-fine-tuned LLMs. Trained with Reinforcement Learning, PILLOW exhibits commensurate performance on various evaluation metrics compared with typical instruction fine-tuning methods, utilizing only consumer-grade GPU resources and exhibiting a large reduction in computational costs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_05621 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | PILLOW: Enhancing Efficient Instruction Fine-tuning via Prompt Matching Qi, Zhenting Tan, Xiaoyu Shi, Shaojie Qu, Chao Xu, Yinghui Qi, Yuan Computation and Language Instruction fine-tuning has conventionally been employed to adapt Large Language Models (LLMs) to a variety of tasks. Nonetheless, this technique often necessitates substantial computational resources, making it impractical for deployment by individuals or small-scale entities. Recently, Low-Rank Adaptation (LoRA) has become a promising alternative, offering high capabilities on par with full tuning with reduced resource overhead. However, attaining satisfactory performance through the fine-tuning of LoRA is a non-trivial challenge. In this paper, we propose PILLOW, which aims to improve LoRA's performance by a discrimination-based prompting method, leveraging LLMs' In-Context Learning ability. PILLOW incorporates a matching network that selects prompts from a user-defined prompt pool, concatenates the selected prompts with the user instruction as input, and performs inference using the LoRA-fine-tuned LLMs. Trained with Reinforcement Learning, PILLOW exhibits commensurate performance on various evaluation metrics compared with typical instruction fine-tuning methods, utilizing only consumer-grade GPU resources and exhibiting a large reduction in computational costs. |
| title | PILLOW: Enhancing Efficient Instruction Fine-tuning via Prompt Matching |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2312.05621 |