How Redundant Is the Transformer Stack in Speech Representation Models?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dorszewski, Teresa, Jacobsen, Albert Kjøller, Tětková, Lenka, Hansen, Lars Kai
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910787519381504
author Dorszewski, Teresa
Jacobsen, Albert Kjøller
Tětková, Lenka
Hansen, Lars Kai
author_facet Dorszewski, Teresa
Jacobsen, Albert Kjøller
Tětková, Lenka
Hansen, Lars Kai
contents Self-supervised speech representation models, particularly those leveraging transformer architectures, have demonstrated remarkable performance across various tasks such as speech recognition, speaker identification, and emotion detection. Recent studies on transformer models revealed a high redundancy between layers and the potential for significant pruning, which we will investigate here for transformer-based speech representation models. We perform a detailed analysis of layer similarity in speech representation models using three similarity metrics: cosine similarity, centered kernel alignment, and mutual nearest-neighbor alignment. Our findings reveal a block-like structure of high similarity, suggesting two main processing steps and significant redundancy of layers. We demonstrate the effectiveness of pruning transformer-based speech representation models without the need for post-training, achieving up to 40% reduction in transformer layers while maintaining over 95% of the model's predictive capacity. Furthermore, we employ a knowledge distillation method to substitute the entire transformer stack with mimicking layers, reducing the network size 95-98% and the inference time by up to 94%. This substantial decrease in computational load occurs without considerable performance loss, suggesting that the transformer stack is almost completely redundant for downstream applications of speech representation models.
format Preprint
id arxiv_https___arxiv_org_abs_2409_16302
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Redundant Is the Transformer Stack in Speech Representation Models?
Dorszewski, Teresa
Jacobsen, Albert Kjøller
Tětková, Lenka
Hansen, Lars Kai
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Self-supervised speech representation models, particularly those leveraging transformer architectures, have demonstrated remarkable performance across various tasks such as speech recognition, speaker identification, and emotion detection. Recent studies on transformer models revealed a high redundancy between layers and the potential for significant pruning, which we will investigate here for transformer-based speech representation models. We perform a detailed analysis of layer similarity in speech representation models using three similarity metrics: cosine similarity, centered kernel alignment, and mutual nearest-neighbor alignment. Our findings reveal a block-like structure of high similarity, suggesting two main processing steps and significant redundancy of layers. We demonstrate the effectiveness of pruning transformer-based speech representation models without the need for post-training, achieving up to 40% reduction in transformer layers while maintaining over 95% of the model's predictive capacity. Furthermore, we employ a knowledge distillation method to substitute the entire transformer stack with mimicking layers, reducing the network size 95-98% and the inference time by up to 94%. This substantial decrease in computational load occurs without considerable performance loss, suggesting that the transformer stack is almost completely redundant for downstream applications of speech representation models.
title How Redundant Is the Transformer Stack in Speech Representation Models?
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2409.16302