Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuhlmann, Michael, Cord-Landwehr, Tobias, Haeb-Umbach, Reinhold
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914583294246912
author Kuhlmann, Michael
Cord-Landwehr, Tobias
Haeb-Umbach, Reinhold
author_facet Kuhlmann, Michael
Cord-Landwehr, Tobias
Haeb-Umbach, Reinhold
contents Automatic subjective speech quality assessment (SSQA) traditionally estimates speech quality on an utterance or system level. While this resolution was adequate for older transmission or synthesis systems that produced speech signals of mediocre quality, modern systems generate high-quality speech with degradations that may occur only locally. With suitable model architectures and regularization losses, SSQA models trained with utterance-level targets can also yield useful local predictions of speech quality. In this work, we extend such models to produce frame-level embeddings that cluster by degradation type. Specifically, we employ a partial mix-up strategy on a parallel corpus of clean and degraded utterances and apply a contrastive loss to distinguish between degradation types. Through experiments on both in- and out-of-domain data, we demonstrate that our approach improves degradation detection and enables the identification of degradation types by analyzing embedding clusters.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21332
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals
Kuhlmann, Michael
Cord-Landwehr, Tobias
Haeb-Umbach, Reinhold
Audio and Speech Processing
Automatic subjective speech quality assessment (SSQA) traditionally estimates speech quality on an utterance or system level. While this resolution was adequate for older transmission or synthesis systems that produced speech signals of mediocre quality, modern systems generate high-quality speech with degradations that may occur only locally. With suitable model architectures and regularization losses, SSQA models trained with utterance-level targets can also yield useful local predictions of speech quality. In this work, we extend such models to produce frame-level embeddings that cluster by degradation type. Specifically, we employ a partial mix-up strategy on a parallel corpus of clean and degraded utterances and apply a contrastive loss to distinguish between degradation types. Through experiments on both in- and out-of-domain data, we demonstrate that our approach improves degradation detection and enables the identification of degradation types by analyzing embedding clusters.
title Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals
topic Audio and Speech Processing
url https://arxiv.org/abs/2605.21332