Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cumlin, Fredrik, Liang, Xinyu, Ghosh, Anubhab, Chatterjee, Saikat
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911098564771840
author Cumlin, Fredrik
Liang, Xinyu
Ghosh, Anubhab
Chatterjee, Saikat
author_facet Cumlin, Fredrik
Liang, Xinyu
Ghosh, Anubhab
Chatterjee, Saikat
contents Non-intrusive speech quality assessment (SQA) systems suffer from limited training data and costly human annotations, hindering their generalization to real-time conferencing calls. In this work, we propose leveraging large language models (LLMs) as pseudo-raters for speech quality to address these data bottlenecks. We construct LibriAugmented, a dataset consisting of 101,129 speech clips with simulated degradations labeled by a fine-tuned auditory LLM (Vicuna-7b-v1.5). We compare three training strategies: using human-labeled data, using LLM-labeled data, and a two-stage approach (pretraining on LLM labels, then fine-tuning on human labels), using both DNSMOS Pro and DeePMOS. We test on several datasets across languages and quality degradations. While LLM-labeled training yields mixed results compared to human-labeled training, we provide empirical evidence that the two-stage approach improves the generalization performance (e.g., DNSMOS Pro achieves 0.63 vs. 0.55 PCC on NISQA_TEST_LIVETALK and 0.73 vs. 0.65 PCC on Tencent with reverb). Our findings demonstrate the potential of using LLMs as scalable pseudo-raters for speech quality assessment, offering a cost-effective solution to the data limitation problem.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06284
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment
Cumlin, Fredrik
Liang, Xinyu
Ghosh, Anubhab
Chatterjee, Saikat
Audio and Speech Processing
Non-intrusive speech quality assessment (SQA) systems suffer from limited training data and costly human annotations, hindering their generalization to real-time conferencing calls. In this work, we propose leveraging large language models (LLMs) as pseudo-raters for speech quality to address these data bottlenecks. We construct LibriAugmented, a dataset consisting of 101,129 speech clips with simulated degradations labeled by a fine-tuned auditory LLM (Vicuna-7b-v1.5). We compare three training strategies: using human-labeled data, using LLM-labeled data, and a two-stage approach (pretraining on LLM labels, then fine-tuning on human labels), using both DNSMOS Pro and DeePMOS. We test on several datasets across languages and quality degradations. While LLM-labeled training yields mixed results compared to human-labeled training, we provide empirical evidence that the two-stage approach improves the generalization performance (e.g., DNSMOS Pro achieves 0.63 vs. 0.55 PCC on NISQA_TEST_LIVETALK and 0.73 vs. 0.65 PCC on Tencent with reverb). Our findings demonstrate the potential of using LLMs as scalable pseudo-raters for speech quality assessment, offering a cost-effective solution to the data limitation problem.
title Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.06284