WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Emon, Jakaria Islam, Alam, Kazi Tamanna, Salek, Md. Abu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916782213693440
author Emon, Jakaria Islam
Alam, Kazi Tamanna
Salek, Md. Abu
author_facet Emon, Jakaria Islam
Alam, Kazi Tamanna
Salek, Md. Abu
contents Mean Opinion Score (MOS) prediction for text to music systems requires evaluating both overall musical quality and text prompt alignment. This paper introduces WhisQ, a multimodal architecture that addresses this dual-assessment challenge through sequence level co-attention and optimal transport regularization. WhisQ employs the Whisper Base pretrained model for temporal audio encoding and Qwen 3, a 0.6B Small Language Model (SLM), for text encoding, with both maintaining sequence structure for fine grained cross-modal modeling. The architecture features specialized prediction pathways: OMQ is predicted from pooled audio embeddings, while TA leverages bidirectional sequence co-attention between audio and text. Sinkhorn optimal transport loss further enforce semantic alignment in the shared embedding space. On the MusicEval Track-1 dataset, WhisQ achieves substantial improvements over the baseline: 7% improvement in Spearman correlation for OMQ and 14% for TA. Ablation studies reveal that optimal transport regularization provides the largest performance gain (10% SRCC improvement), demonstrating the importance of explicit cross-modal alignment for text-to-music evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05899
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
Emon, Jakaria Islam
Alam, Kazi Tamanna
Salek, Md. Abu
Sound
Artificial Intelligence
Audio and Speech Processing
Mean Opinion Score (MOS) prediction for text to music systems requires evaluating both overall musical quality and text prompt alignment. This paper introduces WhisQ, a multimodal architecture that addresses this dual-assessment challenge through sequence level co-attention and optimal transport regularization. WhisQ employs the Whisper Base pretrained model for temporal audio encoding and Qwen 3, a 0.6B Small Language Model (SLM), for text encoding, with both maintaining sequence structure for fine grained cross-modal modeling. The architecture features specialized prediction pathways: OMQ is predicted from pooled audio embeddings, while TA leverages bidirectional sequence co-attention between audio and text. Sinkhorn optimal transport loss further enforce semantic alignment in the shared embedding space. On the MusicEval Track-1 dataset, WhisQ achieves substantial improvements over the baseline: 7% improvement in Spearman correlation for OMQ and 14% for TA. Ablation studies reveal that optimal transport regularization provides the largest performance gain (10% SRCC improvement), demonstrating the importance of explicit cross-modal alignment for text-to-music evaluation.
title WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2506.05899