Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918135067574272 |
|---|---|
| author | Wisnu, Dyah A. M. G. Zezario, Ryandhimas E. Rini, Stefano Wang, Hsin-Min Tsao, Yu |
| author_facet | Wisnu, Dyah A. M. G. Zezario, Ryandhimas E. Rini, Stefano Wang, Hsin-Min Tsao, Yu |
| contents | We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness--for audio generated by text-to-speech (TTS), text-to-audio (TTA), and text-to-music (TTM) systems. A main challenge is the domain shift between natural training data and synthetic evaluation data. To address this, we combine BEATs, a pretrained transformer-based audio representation model, with a multi-branch long short-term memory (LSTM) predictor and use a triplet loss with buffer-based sampling to structure the embedding space by perceptual similarity. Our results show that this improves embedding discriminability and generalization, enabling domain-robust audio quality assessment without synthetic training data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_03292 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings Wisnu, Dyah A. M. G. Zezario, Ryandhimas E. Rini, Stefano Wang, Hsin-Min Tsao, Yu Audio and Speech Processing Machine Learning Sound We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness--for audio generated by text-to-speech (TTS), text-to-audio (TTA), and text-to-music (TTM) systems. A main challenge is the domain shift between natural training data and synthetic evaluation data. To address this, we combine BEATs, a pretrained transformer-based audio representation model, with a multi-branch long short-term memory (LSTM) predictor and use a triplet loss with buffer-based sampling to structure the embedding space by perceptual similarity. Our results show that this improves embedding discriminability and generalization, enabling domain-robust audio quality assessment without synthetic training data. |
| title | Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings |
| topic | Audio and Speech Processing Machine Learning Sound |
| url | https://arxiv.org/abs/2509.03292 |