Introducing voice timbre attribute detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Jinghao, Sheng, Zhengyan, Chen, Liping, Lee, Kong Aik, Ling, Zhen-Hua
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918067227852800
author He, Jinghao
Sheng, Zhengyan
Chen, Liping
Lee, Kong Aik
Ling, Zhen-Hua
author_facet He, Jinghao
Sheng, Zhengyan
Chen, Liping
Lee, Kong Aik
Ling, Zhen-Hua
contents This paper focuses on explaining the timbre conveyed by speech signals and introduces a task termed voice timbre attribute detection (vTAD). In this task, voice timbre is explained with a set of sensory attributes describing its human perception. A pair of speech utterances is processed, and their intensity is compared in a designated timbre descriptor. Moreover, a framework is proposed, which is built upon the speaker embeddings extracted from the speech utterances. The investigation is conducted on the VCTK-RVA dataset. Experimental examinations on the ECAPA-TDNN and FACodec speaker encoders demonstrated that: 1) the ECAPA-TDNN speaker encoder was more capable in the seen scenario, where the testing speakers were included in the training set; 2) the FACodec speaker encoder was superior in the unseen scenario, where the testing speakers were not part of the training, indicating enhanced generalization capability. The VCTK-RVA dataset and open-source code are available on the website https://github.com/vTAD2025-Challenge/vTAD.
format Preprint
id arxiv_https___arxiv_org_abs_2505_09661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Introducing voice timbre attribute detection
He, Jinghao
Sheng, Zhengyan
Chen, Liping
Lee, Kong Aik
Ling, Zhen-Hua
Sound
Artificial Intelligence
Audio and Speech Processing
This paper focuses on explaining the timbre conveyed by speech signals and introduces a task termed voice timbre attribute detection (vTAD). In this task, voice timbre is explained with a set of sensory attributes describing its human perception. A pair of speech utterances is processed, and their intensity is compared in a designated timbre descriptor. Moreover, a framework is proposed, which is built upon the speaker embeddings extracted from the speech utterances. The investigation is conducted on the VCTK-RVA dataset. Experimental examinations on the ECAPA-TDNN and FACodec speaker encoders demonstrated that: 1) the ECAPA-TDNN speaker encoder was more capable in the seen scenario, where the testing speakers were included in the training set; 2) the FACodec speaker encoder was superior in the unseen scenario, where the testing speakers were not part of the training, indicating enhanced generalization capability. The VCTK-RVA dataset and open-source code are available on the website https://github.com/vTAD2025-Challenge/vTAD.
title Introducing voice timbre attribute detection
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.09661