An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915094230728704 |
|---|---|
| author | Li, Jingyu Chiu, Aemon Yat Fei Lee, Tan |
| author_facet | Li, Jingyu Chiu, Aemon Yat Fei Lee, Tan |
| contents | Language mismatch is among the most common and challenging domain mismatches in deploying speaker verification (SV) systems. Adversarial reprogramming has shown promising results in cross-language adaptation for SV. The reprogramming is implemented by padding learnable parameters on the two sides of input speech signals. In this paper, we investigate the relationship between the number of padded parameters and the performance of the reprogrammed models. Sufficient experiments are conducted with different scales of SV models and datasets. The results demonstrate that reprogramming consistently improves the performance of cross-language SV, while the improvement is saturated or even degraded when using larger padding lengths. The performance is mainly determined by the capacity of the original SV models instead of the number of padded parameters. The SV models with larger scales have higher upper bounds in performance and can endure longer padding without performance degradation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_11353 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems Li, Jingyu Chiu, Aemon Yat Fei Lee, Tan Audio and Speech Processing Sound Language mismatch is among the most common and challenging domain mismatches in deploying speaker verification (SV) systems. Adversarial reprogramming has shown promising results in cross-language adaptation for SV. The reprogramming is implemented by padding learnable parameters on the two sides of input speech signals. In this paper, we investigate the relationship between the number of padded parameters and the performance of the reprogrammed models. Sufficient experiments are conducted with different scales of SV models and datasets. The results demonstrate that reprogramming consistently improves the performance of cross-language SV, while the improvement is saturated or even degraded when using larger padding lengths. The performance is mainly determined by the capacity of the original SV models instead of the number of padded parameters. The SV models with larger scales have higher upper bounds in performance and can endure longer padding without performance degradation. |
| title | An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2411.11353 |