Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914474130145280 |
|---|---|
| author | Zezario, Ryandhimas E. Wisnu, Dyah A. M. G. Fu, Szu-Wei Siniscalchi, Sabato Marco Wang, Hsin-Min Tsao, Yu |
| author_facet | Zezario, Ryandhimas E. Wisnu, Dyah A. M. G. Fu, Szu-Wei Siniscalchi, Sabato Marco Wang, Hsin-Min Tsao, Yu |
| contents | In this paper, we introduce GatherMOS, a novel framework that leverages large language models (LLM) as meta-evaluators to aggregate diverse signals into quality predictions. GatherMOS integrates lightweight acoustic descriptors with pseudo-labels from DNSMOS and VQScore, enabling the LLM to reason over heterogeneous inputs and infer perceptual mean opinion scores (MOS). We further explore both zero-shot and few-shot in-context learning setups, showing that zero-shot GatherMOS maintains stable performance across diverse conditions, while few-shot guidance yields large gains when support samples match the test conditions. Experiments on the VoiceBank-DEMAND dataset demonstrate that GatherMOS consistently outperforms DNSMOS, VQScore, naive score averaging, and even learning-based models such as CNN-BLSTM and MOS-SSL when trained under limited labeled-data conditions. These results highlight the potential of LLM-based aggregation as a practical strategy for non-intrusive speech quality evaluation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_13528 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models Zezario, Ryandhimas E. Wisnu, Dyah A. M. G. Fu, Szu-Wei Siniscalchi, Sabato Marco Wang, Hsin-Min Tsao, Yu Audio and Speech Processing Sound In this paper, we introduce GatherMOS, a novel framework that leverages large language models (LLM) as meta-evaluators to aggregate diverse signals into quality predictions. GatherMOS integrates lightweight acoustic descriptors with pseudo-labels from DNSMOS and VQScore, enabling the LLM to reason over heterogeneous inputs and infer perceptual mean opinion scores (MOS). We further explore both zero-shot and few-shot in-context learning setups, showing that zero-shot GatherMOS maintains stable performance across diverse conditions, while few-shot guidance yields large gains when support samples match the test conditions. Experiments on the VoiceBank-DEMAND dataset demonstrate that GatherMOS consistently outperforms DNSMOS, VQScore, naive score averaging, and even learning-based models such as CNN-BLSTM and MOS-SSL when trained under limited labeled-data conditions. These results highlight the potential of LLM-based aggregation as a practical strategy for non-intrusive speech quality evaluation. |
| title | Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2604.13528 |