LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910044457533440 |
|---|---|
| author | Ohmura, Junki Ito, Yuki Tsunoo, Emiru Sekiya, Toshiyuki Kumakura, Toshiyuki |
| author_facet | Ohmura, Junki Ito, Yuki Tsunoo, Emiru Sekiya, Toshiyuki Kumakura, Toshiyuki |
| contents | Numerical voice impression (VI) control (e.g., scaling brightness) enables fine-grained control in text-to-speech (TTS). However, it faces two challenges: no public corpus and impression leakage, where reference audio biases synthesized voice away from the target VI. To address the first challenge, we introduce LibriTTS-VI, the first public VI corpus built on LibriTTS-R. For the second, we hypothesize a single reference causes leakage by entangling speaker identity and VI. To mitigate this, we propose 1) disentangled training with two utterances from the same speaker for speaker and VI conditioning, and 2) a reference-free method controlling the impression solely via target VI. Experimentally, our best method improves controllability: 11-dimensional VI mean squared error drops from 0.61 to 0.41 objectively and 1.15 to 0.92 subjectively. A comparison with a prompt-based TTS reveals imprecise numerical control and entanglement between VI and text semantics, which our methods overcome. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_15626 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control Ohmura, Junki Ito, Yuki Tsunoo, Emiru Sekiya, Toshiyuki Kumakura, Toshiyuki Sound Audio and Speech Processing Numerical voice impression (VI) control (e.g., scaling brightness) enables fine-grained control in text-to-speech (TTS). However, it faces two challenges: no public corpus and impression leakage, where reference audio biases synthesized voice away from the target VI. To address the first challenge, we introduce LibriTTS-VI, the first public VI corpus built on LibriTTS-R. For the second, we hypothesize a single reference causes leakage by entangling speaker identity and VI. To mitigate this, we propose 1) disentangled training with two utterances from the same speaker for speaker and VI conditioning, and 2) a reference-free method controlling the impression solely via target VI. Experimentally, our best method improves controllability: 11-dimensional VI mean squared error drops from 0.61 to 0.41 objectively and 1.15 to 0.92 subjectively. A comparison with a prompt-based TTS reveals imprecise numerical control and entanglement between VI and text semantics, which our methods overcome. |
| title | LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2509.15626 |