Multi-interaction TTS toward professional recording reproduction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kanagawa, Hiroki, Fujita, Kenichi, Watanabe, Aya, Ijima, Yusuke
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913921951072256
author Kanagawa, Hiroki
Fujita, Kenichi
Watanabe, Aya
Ijima, Yusuke
author_facet Kanagawa, Hiroki
Fujita, Kenichi
Watanabe, Aya
Ijima, Yusuke
contents Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has been overlooked in text-to-speech synthesis (TTS). As a result, fine-grained style refinement after the initial synthesis is not possible, even though the synthesized speech often deviates from the user's intended style. To address this issue, we propose a TTS method with multi-step interaction that allows users to intuitively and rapidly refine synthesized speech. Our approach models the interaction between the TTS model and its user to emulate the relationship between voice actors and voice directors. Experiments show that the proposed model with its corresponding dataset enables iterative style refinements in accordance with users' directions, thus demonstrating its multi-interaction capability. Sample audios are available: https://ntt-hilab-gensp.github.io/ssw13multiinteractiontts/
format Preprint
id arxiv_https___arxiv_org_abs_2507_00808
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-interaction TTS toward professional recording reproduction
Kanagawa, Hiroki
Fujita, Kenichi
Watanabe, Aya
Ijima, Yusuke
Sound
Computation and Language
Audio and Speech Processing
Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has been overlooked in text-to-speech synthesis (TTS). As a result, fine-grained style refinement after the initial synthesis is not possible, even though the synthesized speech often deviates from the user's intended style. To address this issue, we propose a TTS method with multi-step interaction that allows users to intuitively and rapidly refine synthesized speech. Our approach models the interaction between the TTS model and its user to emulate the relationship between voice actors and voice directors. Experiments show that the proposed model with its corresponding dataset enables iterative style refinements in accordance with users' directions, thus demonstrating its multi-interaction capability. Sample audios are available: https://ntt-hilab-gensp.github.io/ssw13multiinteractiontts/
title Multi-interaction TTS toward professional recording reproduction
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2507.00808