MambaRate: Speech Quality Assessment Across Different Sampling Rates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kakoulidis, Panos, Alexiou, Iakovi, Oh, Junkwang, Jho, Gunu, Hwang, Inchul, Tsiakoulis, Pirros, Chalamandaris, Aimilios
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913944466096128
author Kakoulidis, Panos
Alexiou, Iakovi
Oh, Junkwang
Jho, Gunu
Hwang, Inchul
Tsiakoulis, Pirros
Chalamandaris, Aimilios
author_facet Kakoulidis, Panos
Alexiou, Iakovi
Oh, Junkwang
Jho, Gunu
Hwang, Inchul
Tsiakoulis, Pirros
Chalamandaris, Aimilios
contents We propose MambaRate, which predicts Mean Opinion Scores (MOS) with limited bias regarding the sampling rate of the waveform under evaluation. It is designed for Track 3 of the AudioMOS Challenge 2025, which focuses on predicting MOS for speech in high sampling frequencies. Our model leverages self-supervised embeddings and selective state space modeling. The target ratings are encoded in a continuous representation via Gaussian radial basis functions (RBF). The results of the challenge were based on the system-level Spearman's Rank Correllation Coefficient (SRCC) metric. An initial MambaRate version (T16 system) outperformed the pre-trained baseline (B03) by ~14% in a few-shot setting without pre-training. T16 ranked fourth out of five in the challenge, differing by ~6% from the winning system. We present additional results on the BVCC dataset as well as ablations with different representations as input, which outperform the initial T16 version.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12090
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MambaRate: Speech Quality Assessment Across Different Sampling Rates
Kakoulidis, Panos
Alexiou, Iakovi
Oh, Junkwang
Jho, Gunu
Hwang, Inchul
Tsiakoulis, Pirros
Chalamandaris, Aimilios
Sound
Audio and Speech Processing
We propose MambaRate, which predicts Mean Opinion Scores (MOS) with limited bias regarding the sampling rate of the waveform under evaluation. It is designed for Track 3 of the AudioMOS Challenge 2025, which focuses on predicting MOS for speech in high sampling frequencies. Our model leverages self-supervised embeddings and selective state space modeling. The target ratings are encoded in a continuous representation via Gaussian radial basis functions (RBF). The results of the challenge were based on the system-level Spearman's Rank Correllation Coefficient (SRCC) metric. An initial MambaRate version (T16 system) outperformed the pre-trained baseline (B03) by ~14% in a few-shot setting without pre-training. T16 ranked fourth out of five in the challenge, differing by ~6% from the winning system. We present additional results on the BVCC dataset as well as ablations with different representations as input, which outperform the initial T16 version.
title MambaRate: Speech Quality Assessment Across Different Sampling Rates
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.12090