SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Yueqian, Fu, Yuzhe, Zhang, Jingyang, Liu, Yudong, Zhang, Jianyi, Sun, Jingwei, Li, Hai "Helen", Chen, Yiran
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908289529282560
author Lin, Yueqian
Fu, Yuzhe
Zhang, Jingyang
Liu, Yudong
Zhang, Jianyi
Sun, Jingwei
Li, Hai "Helen"
Chen, Yiran
author_facet Lin, Yueqian
Fu, Yuzhe
Zhang, Jingyang
Liu, Yudong
Zhang, Jianyi
Sun, Jingwei
Li, Hai "Helen"
Chen, Yiran
contents We introduce Speech Information Retrieval (SIR), a new long-context task for Speech Large Language Models (Speech LLMs), and present SPIRAL, a 1,012-sample benchmark testing models' ability to extract critical details from approximately 90-second spoken inputs. While current Speech LLMs excel at short-form tasks, they struggle with the computational and representational demands of longer audio sequences. To address this limitation, we propose SpeechPrune, a training-free token pruning strategy that uses speech-text similarity and approximated attention scores to efficiently discard irrelevant tokens. In SPIRAL, SpeechPrune achieves accuracy improvements of 29% and up to 47% over the original model and the random pruning model at a pruning rate of 20%, respectively. SpeechPrune can maintain network performance even at a pruning level of 80%. This approach highlights the potential of token-level pruning for efficient and scalable long-form speech understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12009
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
Lin, Yueqian
Fu, Yuzhe
Zhang, Jingyang
Liu, Yudong
Zhang, Jianyi
Sun, Jingwei
Li, Hai "Helen"
Chen, Yiran
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
We introduce Speech Information Retrieval (SIR), a new long-context task for Speech Large Language Models (Speech LLMs), and present SPIRAL, a 1,012-sample benchmark testing models' ability to extract critical details from approximately 90-second spoken inputs. While current Speech LLMs excel at short-form tasks, they struggle with the computational and representational demands of longer audio sequences. To address this limitation, we propose SpeechPrune, a training-free token pruning strategy that uses speech-text similarity and approximated attention scores to efficiently discard irrelevant tokens. In SPIRAL, SpeechPrune achieves accuracy improvements of 29% and up to 47% over the original model and the random pruning model at a pruning rate of 20%, respectively. SpeechPrune can maintain network performance even at a pruning level of 80%. This approach highlights the potential of token-level pruning for efficient and scalable long-form speech understanding.
title SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
url https://arxiv.org/abs/2412.12009