Pay Attention to What You Need
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866929725073522688 |
|---|---|
| author | Gao, Yifei Chen, Shaohong Wang, Lei Dai, Ruiting Zhang, Ziyun Ren, Kerui Wu, Jiaji Cheng, Jun |
| author_facet | Gao, Yifei Chen, Shaohong Wang, Lei Dai, Ruiting Zhang, Ziyun Ren, Kerui Wu, Jiaji Cheng, Jun |
| contents | Although large language models (LLMs) have achieved significant success in natural language processing, they still struggle with long-context comprehension. Traditional approaches to mitigating this issue typically rely on fine-tuning or retraining, which is both resource-intensive and challenging to deploy in lightweight industrial settings. In this paper, we investigate the potential to accomplish this without any additional resources. Through an in-depth study of the attention mechanism in LLMs, we propose a method called Scaled ReAttention (SRA) to strengthen LLMs' ability to interpret and retrieve information by strategically manipulating their attention scores during inference. Through extensive experiments, we demonstrate that integrating SRA significantly boosts LLMs' performance on a variety of downstream tasks, highlighting its practical potential for enhancing language understanding without incurring the overhead of traditional training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2307_13365 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Pay Attention to What You Need Gao, Yifei Chen, Shaohong Wang, Lei Dai, Ruiting Zhang, Ziyun Ren, Kerui Wu, Jiaji Cheng, Jun Computation and Language Artificial Intelligence 68T07, 68T50 Although large language models (LLMs) have achieved significant success in natural language processing, they still struggle with long-context comprehension. Traditional approaches to mitigating this issue typically rely on fine-tuning or retraining, which is both resource-intensive and challenging to deploy in lightweight industrial settings. In this paper, we investigate the potential to accomplish this without any additional resources. Through an in-depth study of the attention mechanism in LLMs, we propose a method called Scaled ReAttention (SRA) to strengthen LLMs' ability to interpret and retrieve information by strategically manipulating their attention scores during inference. Through extensive experiments, we demonstrate that integrating SRA significantly boosts LLMs' performance on a variety of downstream tasks, highlighting its practical potential for enhancing language understanding without incurring the overhead of traditional training. |
| title | Pay Attention to What You Need |
| topic | Computation and Language Artificial Intelligence 68T07, 68T50 |
| url | https://arxiv.org/abs/2307.13365 |