RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kanamori, Yusuke, Okamoto, Yuki, Takano, Taisei, Takamichi, Shinnosuke, Saito, Yuki, Saruwatari, Hiroshi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908428113281024
author Kanamori, Yusuke
Okamoto, Yuki
Takano, Taisei
Takamichi, Shinnosuke
Saito, Yuki
Saruwatari, Hiroshi
author_facet Kanamori, Yusuke
Okamoto, Yuki
Takano, Taisei
Takamichi, Shinnosuke
Saito, Yuki
Saruwatari, Hiroshi
contents In text-to-audio (TTA) research, the relevance between input text and output audio is an important evaluation aspect. Traditionally, it has been evaluated from both subjective and objective perspectives. However, subjective evaluation is costly in terms of money and time, and objective evaluation is unclear regarding the correlation to subjective evaluation scores. In this study, we construct RELATE, an open-sourced dataset that subjectively evaluates the relevance. Also, we benchmark a model for automatically predicting the subjective evaluation score from synthesized audio. Our model outperforms a conventional CLAPScore model, and that trend extends to many sound categories.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23582
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
Kanamori, Yusuke
Okamoto, Yuki
Takano, Taisei
Takamichi, Shinnosuke
Saito, Yuki
Saruwatari, Hiroshi
Sound
Audio and Speech Processing
In text-to-audio (TTA) research, the relevance between input text and output audio is an important evaluation aspect. Traditionally, it has been evaluated from both subjective and objective perspectives. However, subjective evaluation is costly in terms of money and time, and objective evaluation is unclear regarding the correlation to subjective evaluation scores. In this study, we construct RELATE, an open-sourced dataset that subjectively evaluates the relevance. Also, we benchmark a model for automatically predicting the subjective evaluation score from synthesized audio. Our model outperforms a conventional CLAPScore model, and that trend extends to many sound categories.
title RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.23582