Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916812265881600 |
|---|---|
| author | Lin, Tzu-Quan Cheng, Hsi-Chun Lee, Hung-yi Tang, Hao |
| author_facet | Lin, Tzu-Quan Cheng, Hsi-Chun Lee, Hung-yi Tang, Hao |
| contents | In recent years, the impact of self-supervised speech Transformers has extended to speaker-related applications. However, little research has explored how these models encode speaker information. In this work, we address this gap by identifying neurons in the feed-forward layers that are correlated with speaker information. Specifically, we analyze neurons associated with k-means clusters of self-supervised features and i-vectors. Our analysis reveals that these clusters correspond to broad phonetic and gender classes, making them suitable for identifying neurons that represent speakers. By protecting these neurons during pruning, we can significantly preserve performance on speaker-related task, demonstrating their crucial role in encoding speaker information. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_21712 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers Lin, Tzu-Quan Cheng, Hsi-Chun Lee, Hung-yi Tang, Hao Computation and Language Sound Audio and Speech Processing In recent years, the impact of self-supervised speech Transformers has extended to speaker-related applications. However, little research has explored how these models encode speaker information. In this work, we address this gap by identifying neurons in the feed-forward layers that are correlated with speaker information. Specifically, we analyze neurons associated with k-means clusters of self-supervised features and i-vectors. Our analysis reveals that these clusters correspond to broad phonetic and gender classes, making them suitable for identifying neurons that represent speakers. By protecting these neurons during pruning, we can significantly preserve performance on speaker-related task, demonstrating their crucial role in encoding speaker information. |
| title | Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.21712 |