An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Jingyu, Chiu, Aemon Yat Fei, Lee, Tan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915094230728704
author Li, Jingyu
Chiu, Aemon Yat Fei
Lee, Tan
author_facet Li, Jingyu
Chiu, Aemon Yat Fei
Lee, Tan
contents Language mismatch is among the most common and challenging domain mismatches in deploying speaker verification (SV) systems. Adversarial reprogramming has shown promising results in cross-language adaptation for SV. The reprogramming is implemented by padding learnable parameters on the two sides of input speech signals. In this paper, we investigate the relationship between the number of padded parameters and the performance of the reprogrammed models. Sufficient experiments are conducted with different scales of SV models and datasets. The results demonstrate that reprogramming consistently improves the performance of cross-language SV, while the improvement is saturated or even degraded when using larger padding lengths. The performance is mainly determined by the capacity of the original SV models instead of the number of padded parameters. The SV models with larger scales have higher upper bounds in performance and can endure longer padding without performance degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11353
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems
Li, Jingyu
Chiu, Aemon Yat Fei
Lee, Tan
Audio and Speech Processing
Sound
Language mismatch is among the most common and challenging domain mismatches in deploying speaker verification (SV) systems. Adversarial reprogramming has shown promising results in cross-language adaptation for SV. The reprogramming is implemented by padding learnable parameters on the two sides of input speech signals. In this paper, we investigate the relationship between the number of padded parameters and the performance of the reprogrammed models. Sufficient experiments are conducted with different scales of SV models and datasets. The results demonstrate that reprogramming consistently improves the performance of cross-language SV, while the improvement is saturated or even degraded when using larger padding lengths. The performance is mainly determined by the capacity of the original SV models instead of the number of padded parameters. The SV models with larger scales have higher upper bounds in performance and can endure longer padding without performance degradation.
title An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2411.11353