Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2407.15131 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916331517902848 |
|---|---|
| author | Park, Junyoung Kang, Myeonggu Han, Yunki Kim, Yanggon Shin, Jaekang Kim, Lee-Sup |
| author_facet | Park, Junyoung Kang, Myeonggu Han, Yunki Kim, Yanggon Shin, Jaekang Kim, Lee-Sup |
| contents | The attention mechanism in text generation is memory-bounded due to its sequential characteristics. Therefore, off-chip memory accesses should be minimized for faster execution. Although previous methods addressed this by pruning unimportant tokens, they fall short in selectively removing tokens with near-zero attention probabilities in each instance. Our method estimates the probability before the softmax function, effectively removing low probability tokens and achieving an 12.1x pruning ratio without fine-tuning. Additionally, we present a hardware design supporting seamless on-demand off-chip access. Our approach shows 2.6x reduced memory accesses, leading to an average 2.3x speedup and a 2.4x energy efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_15131 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability Estimation Park, Junyoung Kang, Myeonggu Han, Yunki Kim, Yanggon Shin, Jaekang Kim, Lee-Sup Hardware Architecture Machine Learning The attention mechanism in text generation is memory-bounded due to its sequential characteristics. Therefore, off-chip memory accesses should be minimized for faster execution. Although previous methods addressed this by pruning unimportant tokens, they fall short in selectively removing tokens with near-zero attention probabilities in each instance. Our method estimates the probability before the softmax function, effectively removing low probability tokens and achieving an 12.1x pruning ratio without fine-tuning. Additionally, we present a hardware design supporting seamless on-demand off-chip access. Our approach shows 2.6x reduced memory accesses, leading to an average 2.3x speedup and a 2.4x energy efficiency. |
| title | Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability Estimation |
| topic | Hardware Architecture Machine Learning |
| url | https://arxiv.org/abs/2407.15131 |