The AudioMOS Challenge 2025

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Wen-Chin, Wang, Hui, Liu, Cheng, Wu, Yi-Chiao, Tjandra, Andros, Hsu, Wei-Ning, Cooper, Erica, Qin, Yong, Toda, Tomoki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912563710656512
author Huang, Wen-Chin
Wang, Hui
Liu, Cheng
Wu, Yi-Chiao
Tjandra, Andros
Hsu, Wei-Ning
Cooper, Erica
Qin, Yong
Toda, Tomoki
author_facet Huang, Wen-Chin
Wang, Hui
Liu, Cheng
Wu, Yi-Chiao
Tjandra, Andros
Hsu, Wei-Ning
Cooper, Erica
Qin, Yong
Toda, Tomoki
contents This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set consists of text-to-speech, text-to-audio, and text-to-music samples. The third track focuses on synthetic speech quality assessment in different sampling rates. The challenge attracted 24 unique teams from both academia and industry, and improvements over the baselines were confirmed. The outcome of this challenge is expected to facilitate development and progress in the field of automatic evaluation for audio generation systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_01336
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The AudioMOS Challenge 2025
Huang, Wen-Chin
Wang, Hui
Liu, Cheng
Wu, Yi-Chiao
Tjandra, Andros
Hsu, Wei-Ning
Cooper, Erica
Qin, Yong
Toda, Tomoki
Sound
Audio and Speech Processing
This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set consists of text-to-speech, text-to-audio, and text-to-music samples. The third track focuses on synthetic speech quality assessment in different sampling rates. The challenge attracted 24 unique teams from both academia and industry, and improvements over the baselines were confirmed. The outcome of this challenge is expected to facilitate development and progress in the field of automatic evaluation for audio generation systems.
title The AudioMOS Challenge 2025
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.01336