RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908428113281024 |
|---|---|
| author | Kanamori, Yusuke Okamoto, Yuki Takano, Taisei Takamichi, Shinnosuke Saito, Yuki Saruwatari, Hiroshi |
| author_facet | Kanamori, Yusuke Okamoto, Yuki Takano, Taisei Takamichi, Shinnosuke Saito, Yuki Saruwatari, Hiroshi |
| contents | In text-to-audio (TTA) research, the relevance between input text and output audio is an important evaluation aspect. Traditionally, it has been evaluated from both subjective and objective perspectives. However, subjective evaluation is costly in terms of money and time, and objective evaluation is unclear regarding the correlation to subjective evaluation scores. In this study, we construct RELATE, an open-sourced dataset that subjectively evaluates the relevance. Also, we benchmark a model for automatically predicting the subjective evaluation score from synthesized audio. Our model outperforms a conventional CLAPScore model, and that trend extends to many sound categories. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_23582 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Kanamori, Yusuke Okamoto, Yuki Takano, Taisei Takamichi, Shinnosuke Saito, Yuki Saruwatari, Hiroshi Sound Audio and Speech Processing In text-to-audio (TTA) research, the relevance between input text and output audio is an important evaluation aspect. Traditionally, it has been evaluated from both subjective and objective perspectives. However, subjective evaluation is costly in terms of money and time, and objective evaluation is unclear regarding the correlation to subjective evaluation scores. In this study, we construct RELATE, an open-sourced dataset that subjectively evaluates the relevance. Also, we benchmark a model for automatically predicting the subjective evaluation score from synthesized audio. Our model outperforms a conventional CLAPScore model, and that trend extends to many sound categories. |
| title | RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.23582 |