Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wisnu, Dyah A. M. G., Zezario, Ryandhimas E., Rini, Stefano, Wang, Hsin-Min, Tsao, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918135067574272
author Wisnu, Dyah A. M. G.
Zezario, Ryandhimas E.
Rini, Stefano
Wang, Hsin-Min
Tsao, Yu
author_facet Wisnu, Dyah A. M. G.
Zezario, Ryandhimas E.
Rini, Stefano
Wang, Hsin-Min
Tsao, Yu
contents We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness--for audio generated by text-to-speech (TTS), text-to-audio (TTA), and text-to-music (TTM) systems. A main challenge is the domain shift between natural training data and synthetic evaluation data. To address this, we combine BEATs, a pretrained transformer-based audio representation model, with a multi-branch long short-term memory (LSTM) predictor and use a triplet loss with buffer-based sampling to structure the embedding space by perceptual similarity. Our results show that this improves embedding discriminability and generalization, enabling domain-robust audio quality assessment without synthetic training data.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03292
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
Wisnu, Dyah A. M. G.
Zezario, Ryandhimas E.
Rini, Stefano
Wang, Hsin-Min
Tsao, Yu
Audio and Speech Processing
Machine Learning
Sound
We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness--for audio generated by text-to-speech (TTS), text-to-audio (TTA), and text-to-music (TTM) systems. A main challenge is the domain shift between natural training data and synthetic evaluation data. To address this, we combine BEATs, a pretrained transformer-based audio representation model, with a multi-branch long short-term memory (LSTM) predictor and use a triplet loss with buffer-based sampling to structure the embedding space by perceptual similarity. Our results show that this improves embedding discriminability and generalization, enabling domain-robust audio quality assessment without synthetic training data.
title Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2509.03292