Pay Attention to What You Need

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Yifei, Chen, Shaohong, Wang, Lei, Dai, Ruiting, Zhang, Ziyun, Ren, Kerui, Wu, Jiaji, Cheng, Jun
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929725073522688
author Gao, Yifei
Chen, Shaohong
Wang, Lei
Dai, Ruiting
Zhang, Ziyun
Ren, Kerui
Wu, Jiaji
Cheng, Jun
author_facet Gao, Yifei
Chen, Shaohong
Wang, Lei
Dai, Ruiting
Zhang, Ziyun
Ren, Kerui
Wu, Jiaji
Cheng, Jun
contents Although large language models (LLMs) have achieved significant success in natural language processing, they still struggle with long-context comprehension. Traditional approaches to mitigating this issue typically rely on fine-tuning or retraining, which is both resource-intensive and challenging to deploy in lightweight industrial settings. In this paper, we investigate the potential to accomplish this without any additional resources. Through an in-depth study of the attention mechanism in LLMs, we propose a method called Scaled ReAttention (SRA) to strengthen LLMs' ability to interpret and retrieve information by strategically manipulating their attention scores during inference. Through extensive experiments, we demonstrate that integrating SRA significantly boosts LLMs' performance on a variety of downstream tasks, highlighting its practical potential for enhancing language understanding without incurring the overhead of traditional training.
format Preprint
id arxiv_https___arxiv_org_abs_2307_13365
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Pay Attention to What You Need
Gao, Yifei
Chen, Shaohong
Wang, Lei
Dai, Ruiting
Zhang, Ziyun
Ren, Kerui
Wu, Jiaji
Cheng, Jun
Computation and Language
Artificial Intelligence
68T07, 68T50
Although large language models (LLMs) have achieved significant success in natural language processing, they still struggle with long-context comprehension. Traditional approaches to mitigating this issue typically rely on fine-tuning or retraining, which is both resource-intensive and challenging to deploy in lightweight industrial settings. In this paper, we investigate the potential to accomplish this without any additional resources. Through an in-depth study of the attention mechanism in LLMs, we propose a method called Scaled ReAttention (SRA) to strengthen LLMs' ability to interpret and retrieve information by strategically manipulating their attention scores during inference. Through extensive experiments, we demonstrate that integrating SRA significantly boosts LLMs' performance on a variety of downstream tasks, highlighting its practical potential for enhancing language understanding without incurring the overhead of traditional training.
title Pay Attention to What You Need
topic Computation and Language
Artificial Intelligence
68T07, 68T50
url https://arxiv.org/abs/2307.13365