InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Kexin, Tu, Qian, Fan, Liwei, Yang, Chenchen, Zhang, Dong, Li, Shimin, Fei, Zhaoye, Cheng, Qinyuan, Qiu, Xipeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912441823133696
author Huang, Kexin
Tu, Qian
Fan, Liwei
Yang, Chenchen
Zhang, Dong
Li, Shimin
Fei, Zhaoye
Cheng, Qinyuan
Qiu, Xipeng
author_facet Huang, Kexin
Tu, Qian
Fan, Liwei
Yang, Chenchen
Zhang, Dong
Li, Shimin
Fei, Zhaoye
Cheng, Qinyuan
Qiu, Xipeng
contents In modern speech synthesis, paralinguistic information--such as a speaker's vocal timbre, emotional state, and dynamic prosody--plays a critical role in conveying nuance beyond mere semantics. Traditional Text-to-Speech (TTS) systems rely on fixed style labels or inserting a speech prompt to control these cues, which severely limits flexibility. Recent attempts seek to employ natural-language instructions to modulate paralinguistic features, substantially improving the generalization of instruction-driven TTS models. Although many TTS systems now support customized synthesis via textual description, their actual ability to interpret and execute complex instructions remains largely unexplored. In addition, there is still a shortage of high-quality benchmarks and automated evaluation metrics specifically designed for instruction-based TTS, which hinders accurate assessment and iterative optimization of these models. To address these limitations, we introduce InstructTTSEval, a benchmark for measuring the capability of complex natural-language style control. We introduce three tasks, namely Acoustic-Parameter Specification, Descriptive-Style Directive, and Role-Play, including English and Chinese subsets, each with 1k test cases (6k in total) paired with reference audio. We leverage Gemini as an automatic judge to assess their instruction-following abilities. Our evaluation of accessible instruction-following TTS systems highlights substantial room for further improvement. We anticipate that InstructTTSEval will drive progress toward more powerful, flexible, and accurate instruction-following TTS.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16381
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
Huang, Kexin
Tu, Qian
Fan, Liwei
Yang, Chenchen
Zhang, Dong
Li, Shimin
Fei, Zhaoye
Cheng, Qinyuan
Qiu, Xipeng
Computation and Language
Sound
Audio and Speech Processing
In modern speech synthesis, paralinguistic information--such as a speaker's vocal timbre, emotional state, and dynamic prosody--plays a critical role in conveying nuance beyond mere semantics. Traditional Text-to-Speech (TTS) systems rely on fixed style labels or inserting a speech prompt to control these cues, which severely limits flexibility. Recent attempts seek to employ natural-language instructions to modulate paralinguistic features, substantially improving the generalization of instruction-driven TTS models. Although many TTS systems now support customized synthesis via textual description, their actual ability to interpret and execute complex instructions remains largely unexplored. In addition, there is still a shortage of high-quality benchmarks and automated evaluation metrics specifically designed for instruction-based TTS, which hinders accurate assessment and iterative optimization of these models. To address these limitations, we introduce InstructTTSEval, a benchmark for measuring the capability of complex natural-language style control. We introduce three tasks, namely Acoustic-Parameter Specification, Descriptive-Style Directive, and Role-Play, including English and Chinese subsets, each with 1k test cases (6k in total) paired with reference audio. We leverage Gemini as an automatic judge to assess their instruction-following abilities. Our evaluation of accessible instruction-following TTS systems highlights substantial room for further improvement. We anticipate that InstructTTSEval will drive progress toward more powerful, flexible, and accurate instruction-following TTS.
title InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.16381