TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Hui, Liu, Cheng, Chen, Junyang, Liu, Haoze, Jia, Yuhang, Zhao, Shiwan, Zhou, Jiaming, Sun, Haoqin, Bu, Hui, Qin, Yong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909766677168128
author Wang, Hui
Liu, Cheng
Chen, Junyang
Liu, Haoze
Jia, Yuhang
Zhao, Shiwan
Zhou, Jiaming
Sun, Haoqin
Bu, Hui
Qin, Yong
author_facet Wang, Hui
Liu, Cheng
Chen, Junyang
Liu, Haoze
Jia, Yuhang
Zhao, Shiwan
Zhou, Jiaming
Sun, Haoqin
Bu, Hui
Qin, Yong
contents Text-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generalization, and ethical concerns. We present TTA-Bench, a comprehensive benchmark for evaluating TTA models across functional performance, reliability, and social responsibility. It covers seven dimensions including accuracy, robustness, fairness, and toxicity, and includes 2,999 diverse prompts generated through automated and manual methods. We introduce a unified evaluation protocol that combines objective metrics with over 118,000 human annotations from both experts and general users. Ten state-of-the-art models are benchmarked under this framework, offering detailed insights into their strengths and limitations. TTA-Bench establishes a new standard for holistic and responsible evaluation of TTA systems. The dataset and evaluation tools are open-sourced at https://nku-hlt.github.io/tta-bench/.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02398
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
Wang, Hui
Liu, Cheng
Chen, Junyang
Liu, Haoze
Jia, Yuhang
Zhao, Shiwan
Zhou, Jiaming
Sun, Haoqin
Bu, Hui
Qin, Yong
Sound
Audio and Speech Processing
Text-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generalization, and ethical concerns. We present TTA-Bench, a comprehensive benchmark for evaluating TTA models across functional performance, reliability, and social responsibility. It covers seven dimensions including accuracy, robustness, fairness, and toxicity, and includes 2,999 diverse prompts generated through automated and manual methods. We introduce a unified evaluation protocol that combines objective metrics with over 118,000 human annotations from both experts and general users. Ten state-of-the-art models are benchmarked under this framework, offering detailed insights into their strengths and limitations. TTA-Bench establishes a new standard for holistic and responsible evaluation of TTA systems. The dataset and evaluation tools are open-sourced at https://nku-hlt.github.io/tta-bench/.
title TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.02398