Salvato in:
Dettagli Bibliografici
Autori principali: Park, Junyoung, Kang, Myeonggu, Han, Yunki, Kim, Yanggon, Shin, Jaekang, Kim, Lee-Sup
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2407.15131
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916331517902848
author Park, Junyoung
Kang, Myeonggu
Han, Yunki
Kim, Yanggon
Shin, Jaekang
Kim, Lee-Sup
author_facet Park, Junyoung
Kang, Myeonggu
Han, Yunki
Kim, Yanggon
Shin, Jaekang
Kim, Lee-Sup
contents The attention mechanism in text generation is memory-bounded due to its sequential characteristics. Therefore, off-chip memory accesses should be minimized for faster execution. Although previous methods addressed this by pruning unimportant tokens, they fall short in selectively removing tokens with near-zero attention probabilities in each instance. Our method estimates the probability before the softmax function, effectively removing low probability tokens and achieving an 12.1x pruning ratio without fine-tuning. Additionally, we present a hardware design supporting seamless on-demand off-chip access. Our approach shows 2.6x reduced memory accesses, leading to an average 2.3x speedup and a 2.4x energy efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15131
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability Estimation
Park, Junyoung
Kang, Myeonggu
Han, Yunki
Kim, Yanggon
Shin, Jaekang
Kim, Lee-Sup
Hardware Architecture
Machine Learning
The attention mechanism in text generation is memory-bounded due to its sequential characteristics. Therefore, off-chip memory accesses should be minimized for faster execution. Although previous methods addressed this by pruning unimportant tokens, they fall short in selectively removing tokens with near-zero attention probabilities in each instance. Our method estimates the probability before the softmax function, effectively removing low probability tokens and achieving an 12.1x pruning ratio without fine-tuning. Additionally, we present a hardware design supporting seamless on-demand off-chip access. Our approach shows 2.6x reduced memory accesses, leading to an average 2.3x speedup and a 2.4x energy efficiency.
title Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability Estimation
topic Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2407.15131