Voice Impression Control in Zero-Shot TTS

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fujita, Kenichi, Horiguchi, Shota, Ijima, Yusuke
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908838459867136
author Fujita, Kenichi
Horiguchi, Shota
Ijima, Yusuke
author_facet Fujita, Kenichi
Horiguchi, Shota
Ijima, Yusuke
contents Para-/non-linguistic information in speech is pivotal in shaping the listeners' impression. Although zero-shot text-to-speech (TTS) has achieved high speaker fidelity, modulating subtle para-/non-linguistic information to control perceived voice characteristics, i.e., impressions, remains challenging. We have therefore developed a voice impression control method in zero-shot TTS that utilizes a low-dimensional vector to represent the intensities of various voice impression pairs (e.g., dark-bright). The results of both objective and subjective evaluations have demonstrated our method's effectiveness in impression control. Furthermore, generating this vector via a large language model enables target-impression generation from a natural language description of the desired impression, thus eliminating the need for manual optimization. Audio examples are available on our demo page (https://ntt-hilab-gensp.github.io/is2025voiceimpression/).
format Preprint
id arxiv_https___arxiv_org_abs_2506_05688
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Voice Impression Control in Zero-Shot TTS
Fujita, Kenichi
Horiguchi, Shota
Ijima, Yusuke
Sound
Computation and Language
Machine Learning
Audio and Speech Processing
Para-/non-linguistic information in speech is pivotal in shaping the listeners' impression. Although zero-shot text-to-speech (TTS) has achieved high speaker fidelity, modulating subtle para-/non-linguistic information to control perceived voice characteristics, i.e., impressions, remains challenging. We have therefore developed a voice impression control method in zero-shot TTS that utilizes a low-dimensional vector to represent the intensities of various voice impression pairs (e.g., dark-bright). The results of both objective and subjective evaluations have demonstrated our method's effectiveness in impression control. Furthermore, generating this vector via a large language model enables target-impression generation from a natural language description of the desired impression, thus eliminating the need for manual optimization. Audio examples are available on our demo page (https://ntt-hilab-gensp.github.io/is2025voiceimpression/).
title Voice Impression Control in Zero-Shot TTS
topic Sound
Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2506.05688