ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Frummer, Ari, Wang, Helin, Cao, Tianyu, Arbel, Adi, Sieradzki, Yuval, Gal, Oren, Villalba, Jesús, Thebaud, Thomas, Dehak, Najim
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917044483522560
author Frummer, Ari
Wang, Helin
Cao, Tianyu
Arbel, Adi
Sieradzki, Yuval
Gal, Oren
Villalba, Jesús
Thebaud, Thomas
Dehak, Najim
author_facet Frummer, Ari
Wang, Helin
Cao, Tianyu
Arbel, Adi
Sieradzki, Yuval
Gal, Oren
Villalba, Jesús
Thebaud, Thomas
Dehak, Najim
contents Source separation is a crucial pre-processing step for various speech processing tasks, such as automatic speech recognition (ASR). Traditionally, the evaluation metrics for speech separation rely on the matched reference audios and corresponding transcriptions to assess audio quality and intelligibility. However, they cannot be used to evaluate real-world mixtures for which no reference exists. This paper introduces a text-free reference-free evaluation framework based on self-supervised learning (SSL) representations. The proposed framework utilize the mixture and separated tracks to predict jointly audio quality, through the Scale Invariant Signal to Noise Ratio (SI-SNR) metric, and speech intelligibility through the Word Error Rate (WER) metric. We conducted experiments on the WHAMR! dataset, which shows a WER estimation with a mean absolute error (MAE) of 17% and a Pearson correlation coefficient (PCC) of 0.77; and SI-SNR estimation with an MAE of 1.38 and PCC of 0.95. We further demonstrate the robustness of our estimator by using various SSL representations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21014
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring
Frummer, Ari
Wang, Helin
Cao, Tianyu
Arbel, Adi
Sieradzki, Yuval
Gal, Oren
Villalba, Jesús
Thebaud, Thomas
Dehak, Najim
Audio and Speech Processing
Sound
Source separation is a crucial pre-processing step for various speech processing tasks, such as automatic speech recognition (ASR). Traditionally, the evaluation metrics for speech separation rely on the matched reference audios and corresponding transcriptions to assess audio quality and intelligibility. However, they cannot be used to evaluate real-world mixtures for which no reference exists. This paper introduces a text-free reference-free evaluation framework based on self-supervised learning (SSL) representations. The proposed framework utilize the mixture and separated tracks to predict jointly audio quality, through the Scale Invariant Signal to Noise Ratio (SI-SNR) metric, and speech intelligibility through the Word Error Rate (WER) metric. We conducted experiments on the WHAMR! dataset, which shows a WER estimation with a mean absolute error (MAE) of 17% and a Pearson correlation coefficient (PCC) of 0.77; and SI-SNR estimation with an MAE of 1.38 and PCC of 0.95. We further demonstrate the robustness of our estimator by using various SSL representations.
title ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2510.21014