Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ruggiero, Giuseppe, Testa, Matteo, Van de Walle, Jurgen, Di Caro, Luigi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909622449733632
author Ruggiero, Giuseppe
Testa, Matteo
Van de Walle, Jurgen
Di Caro, Luigi
author_facet Ruggiero, Giuseppe
Testa, Matteo
Van de Walle, Jurgen
Di Caro, Luigi
contents Self-supervised learning (SSL) has reduced the reliance on expensive labeling in speech technologies by learning meaningful representations from unannotated data. Since most SSL-based downstream tasks prioritize content information in speech, ideal representations should disentangle content from unwanted variations like speaker characteristics in the SSL representations. However, removing speaker information often degrades other speech components, and existing methods either fail to fully disentangle speaker identity or require resource-intensive models. In this paper, we propose a novel disentanglement method that linearly decomposes SSL representations into speaker-specific and speaker-independent components, effectively generating speaker disentangled representations. Comprehensive experiments show that our approach achieves speaker independence and as such, when applied to content-driven tasks such as voice conversion, our representations yield significant improvements over state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19273
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
Ruggiero, Giuseppe
Testa, Matteo
Van de Walle, Jurgen
Di Caro, Luigi
Sound
Artificial Intelligence
Audio and Speech Processing
Self-supervised learning (SSL) has reduced the reliance on expensive labeling in speech technologies by learning meaningful representations from unannotated data. Since most SSL-based downstream tasks prioritize content information in speech, ideal representations should disentangle content from unwanted variations like speaker characteristics in the SSL representations. However, removing speaker information often degrades other speech components, and existing methods either fail to fully disentangle speaker identity or require resource-intensive models. In this paper, we propose a novel disentanglement method that linearly decomposes SSL representations into speaker-specific and speaker-independent components, effectively generating speaker disentangled representations. Comprehensive experiments show that our approach achieves speaker independence and as such, when applied to content-driven tasks such as voice conversion, our representations yield significant improvements over state-of-the-art methods.
title Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.19273