MMMOS: Multi-domain Multi-axis Audio Quality Assessment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Yi-Cheng, Chen, Jia-Hung, Lee, Hung-yi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909986217525248
author Lin, Yi-Cheng
Chen, Jia-Hung
Lee, Hung-yi
author_facet Lin, Yi-Cheng
Chen, Jia-Hung
Lee, Hung-yi
contents Accurate audio quality estimation is essential for developing and evaluating audio generation, retrieval, and enhancement systems. Existing non-intrusive assessment models predict a single Mean Opinion Score (MOS) for speech, merging diverse perceptual factors and failing to generalize beyond speech. We propose MMMOS, a no-reference, multi-domain audio quality assessment system that estimates four orthogonal axes: Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness across speech, music, and environmental sounds. MMMOS fuses frame-level embeddings from three pretrained encoders (WavLM, MuQ, and M2D) and evaluates three aggregation strategies with four loss functions. By ensembling the top eight models, MMMOS shows a 20-30% reduction in mean squared error and a 4-5% increase in Kendall's τ versus baseline, gains first place in six of eight Production Complexity metrics, and ranks among the top three on 17 of 32 challenge metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2507_04094
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMMOS: Multi-domain Multi-axis Audio Quality Assessment
Lin, Yi-Cheng
Chen, Jia-Hung
Lee, Hung-yi
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Accurate audio quality estimation is essential for developing and evaluating audio generation, retrieval, and enhancement systems. Existing non-intrusive assessment models predict a single Mean Opinion Score (MOS) for speech, merging diverse perceptual factors and failing to generalize beyond speech. We propose MMMOS, a no-reference, multi-domain audio quality assessment system that estimates four orthogonal axes: Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness across speech, music, and environmental sounds. MMMOS fuses frame-level embeddings from three pretrained encoders (WavLM, MuQ, and M2D) and evaluates three aggregation strategies with four loss functions. By ensembling the top eight models, MMMOS shows a 20-30% reduction in mean squared error and a 4-5% increase in Kendall's τ versus baseline, gains first place in six of eight Production Complexity metrics, and ranks among the top three on 17 of 32 challenge metrics.
title MMMOS: Multi-domain Multi-axis Audio Quality Assessment
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.04094