Mitigating Language Mismatch in SSL-Based Speaker Anonymization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhe, Huang, Wen-Chin, Wang, Xin, Miao, Xiaoxiao, Yamagishi, Junichi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915366909771776
author Zhang, Zhe
Huang, Wen-Chin
Wang, Xin
Miao, Xiaoxiao
Yamagishi, Junichi
author_facet Zhang, Zhe
Huang, Wen-Chin
Wang, Xin
Miao, Xiaoxiao
Yamagishi, Junichi
contents Speaker anonymization aims to protect speaker identity while preserving content information and the intelligibility of speech. However, most speaker anonymization systems (SASs) are developed and evaluated using only English, resulting in degraded utility for other languages. This paper investigates language mismatch in SASs for Japanese and Mandarin speech. First, we fine-tune a self-supervised learning (SSL)-based content encoder with Japanese speech to verify effective language adaptation. Then, we propose fine-tuning a multilingual SSL model with Japanese speech and evaluating the SAS in Japanese and Mandarin. Downstream experiments show that fine-tuning an English-only SSL model with the target language enhances intelligibility while maintaining privacy and that multilingual SSL further extends SASs' utility across different languages. These findings highlight the importance of language adaptation and multilingual pre-training of SSLs for robust multilingual speaker anonymization.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00458
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Language Mismatch in SSL-Based Speaker Anonymization
Zhang, Zhe
Huang, Wen-Chin
Wang, Xin
Miao, Xiaoxiao
Yamagishi, Junichi
Audio and Speech Processing
Sound
Speaker anonymization aims to protect speaker identity while preserving content information and the intelligibility of speech. However, most speaker anonymization systems (SASs) are developed and evaluated using only English, resulting in degraded utility for other languages. This paper investigates language mismatch in SASs for Japanese and Mandarin speech. First, we fine-tune a self-supervised learning (SSL)-based content encoder with Japanese speech to verify effective language adaptation. Then, we propose fine-tuning a multilingual SSL model with Japanese speech and evaluating the SAS in Japanese and Mandarin. Downstream experiments show that fine-tuning an English-only SSL model with the target language enhances intelligibility while maintaining privacy and that multilingual SSL further extends SASs' utility across different languages. These findings highlight the importance of language adaptation and multilingual pre-training of SSLs for robust multilingual speaker anonymization.
title Mitigating Language Mismatch in SSL-Based Speaker Anonymization
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2507.00458