LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ohmura, Junki, Ito, Yuki, Tsunoo, Emiru, Sekiya, Toshiyuki, Kumakura, Toshiyuki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910044457533440
author Ohmura, Junki
Ito, Yuki
Tsunoo, Emiru
Sekiya, Toshiyuki
Kumakura, Toshiyuki
author_facet Ohmura, Junki
Ito, Yuki
Tsunoo, Emiru
Sekiya, Toshiyuki
Kumakura, Toshiyuki
contents Numerical voice impression (VI) control (e.g., scaling brightness) enables fine-grained control in text-to-speech (TTS). However, it faces two challenges: no public corpus and impression leakage, where reference audio biases synthesized voice away from the target VI. To address the first challenge, we introduce LibriTTS-VI, the first public VI corpus built on LibriTTS-R. For the second, we hypothesize a single reference causes leakage by entangling speaker identity and VI. To mitigate this, we propose 1) disentangled training with two utterances from the same speaker for speaker and VI conditioning, and 2) a reference-free method controlling the impression solely via target VI. Experimentally, our best method improves controllability: 11-dimensional VI mean squared error drops from 0.61 to 0.41 objectively and 1.15 to 0.92 subjectively. A comparison with a prompt-based TTS reveals imprecise numerical control and entanglement between VI and text semantics, which our methods overcome.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15626
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
Ohmura, Junki
Ito, Yuki
Tsunoo, Emiru
Sekiya, Toshiyuki
Kumakura, Toshiyuki
Sound
Audio and Speech Processing
Numerical voice impression (VI) control (e.g., scaling brightness) enables fine-grained control in text-to-speech (TTS). However, it faces two challenges: no public corpus and impression leakage, where reference audio biases synthesized voice away from the target VI. To address the first challenge, we introduce LibriTTS-VI, the first public VI corpus built on LibriTTS-R. For the second, we hypothesize a single reference causes leakage by entangling speaker identity and VI. To mitigate this, we propose 1) disentangled training with two utterances from the same speaker for speaker and VI conditioning, and 2) a reference-free method controlling the impression solely via target VI. Experimentally, our best method improves controllability: 11-dimensional VI mean squared error drops from 0.61 to 0.41 objectively and 1.15 to 0.92 subjectively. A comparison with a prompt-based TTS reveals imprecise numerical control and entanglement between VI and text semantics, which our methods overcome.
title LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.15626