Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Tzu-Quan, Cheng, Hsi-Chun, Lee, Hung-yi, Tang, Hao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916812265881600
author Lin, Tzu-Quan
Cheng, Hsi-Chun
Lee, Hung-yi
Tang, Hao
author_facet Lin, Tzu-Quan
Cheng, Hsi-Chun
Lee, Hung-yi
Tang, Hao
contents In recent years, the impact of self-supervised speech Transformers has extended to speaker-related applications. However, little research has explored how these models encode speaker information. In this work, we address this gap by identifying neurons in the feed-forward layers that are correlated with speaker information. Specifically, we analyze neurons associated with k-means clusters of self-supervised features and i-vectors. Our analysis reveals that these clusters correspond to broad phonetic and gender classes, making them suitable for identifying neurons that represent speakers. By protecting these neurons during pruning, we can significantly preserve performance on speaker-related task, demonstrating their crucial role in encoding speaker information.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21712
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
Lin, Tzu-Quan
Cheng, Hsi-Chun
Lee, Hung-yi
Tang, Hao
Computation and Language
Sound
Audio and Speech Processing
In recent years, the impact of self-supervised speech Transformers has extended to speaker-related applications. However, little research has explored how these models encode speaker information. In this work, we address this gap by identifying neurons in the feed-forward layers that are correlated with speaker information. Specifically, we analyze neurons associated with k-means clusters of self-supervised features and i-vectors. Our analysis reveals that these clusters correspond to broad phonetic and gender classes, making them suitable for identifying neurons that represent speakers. By protecting these neurons during pruning, we can significantly preserve performance on speaker-related task, demonstrating their crucial role in encoding speaker information.
title Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.21712