Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917997871890432 |
|---|---|
| author | Tomashenko, Natalia Vincent, Emmanuel Tommasi, Marc |
| author_facet | Tomashenko, Natalia Vincent, Emmanuel Tommasi, Marc |
| contents | In this paper, we investigate the impact of speech temporal dynamics in application to automatic speaker verification and speaker voice anonymization tasks. We propose several metrics to perform automatic speaker verification based only on phoneme durations. Experimental results demonstrate that phoneme durations leak some speaker information and can reveal speaker identity from both original and anonymized speech. Thus, this work emphasizes the importance of taking into account the speaker's speech rate and, more importantly, the speaker's phonetic duration characteristics, as well as the need to modify them in order to develop anonymization systems with strong privacy protection capacity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_17164 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization Tomashenko, Natalia Vincent, Emmanuel Tommasi, Marc Audio and Speech Processing Computation and Language Sound In this paper, we investigate the impact of speech temporal dynamics in application to automatic speaker verification and speaker voice anonymization tasks. We propose several metrics to perform automatic speaker verification based only on phoneme durations. Experimental results demonstrate that phoneme durations leak some speaker information and can reveal speaker identity from both original and anonymized speech. Thus, this work emphasizes the importance of taking into account the speaker's speech rate and, more importantly, the speaker's phonetic duration characteristics, as well as the need to modify them in order to develop anonymization systems with strong privacy protection capacity. |
| title | Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization |
| topic | Audio and Speech Processing Computation and Language Sound |
| url | https://arxiv.org/abs/2412.17164 |