The AudioMOS Challenge 2025
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912563710656512 |
|---|---|
| author | Huang, Wen-Chin Wang, Hui Liu, Cheng Wu, Yi-Chiao Tjandra, Andros Hsu, Wei-Ning Cooper, Erica Qin, Yong Toda, Tomoki |
| author_facet | Huang, Wen-Chin Wang, Hui Liu, Cheng Wu, Yi-Chiao Tjandra, Andros Hsu, Wei-Ning Cooper, Erica Qin, Yong Toda, Tomoki |
| contents | This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set consists of text-to-speech, text-to-audio, and text-to-music samples. The third track focuses on synthetic speech quality assessment in different sampling rates. The challenge attracted 24 unique teams from both academia and industry, and improvements over the baselines were confirmed. The outcome of this challenge is expected to facilitate development and progress in the field of automatic evaluation for audio generation systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_01336 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | The AudioMOS Challenge 2025 Huang, Wen-Chin Wang, Hui Liu, Cheng Wu, Yi-Chiao Tjandra, Andros Hsu, Wei-Ning Cooper, Erica Qin, Yong Toda, Tomoki Sound Audio and Speech Processing This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set consists of text-to-speech, text-to-audio, and text-to-music samples. The third track focuses on synthetic speech quality assessment in different sampling rates. The challenge attracted 24 unique teams from both academia and industry, and improvements over the baselines were confirmed. The outcome of this challenge is expected to facilitate development and progress in the field of automatic evaluation for audio generation systems. |
| title | The AudioMOS Challenge 2025 |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2509.01336 |