Reference-aware SFM layers for intrusive intelligibility prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Hanlin, Zhou, Haoshuai, Cao, Boxuan, Mo, Changgeng, Li, Linkai, Wang, Shan X.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918145468399616
author Yu, Hanlin
Zhou, Haoshuai
Cao, Boxuan
Mo, Changgeng
Li, Linkai
Wang, Shan X.
author_facet Yu, Hanlin
Zhou, Haoshuai
Cao, Boxuan
Mo, Changgeng
Li, Linkai
Wang, Shan X.
contents Intrusive speech-intelligibility predictors that exploit explicit reference signals are now widespread, yet they have not consistently surpassed non-intrusive systems. We argue that a primary cause is the limited exploitation of speech foundation models (SFMs). This work revisits intrusive prediction by combining reference conditioning with multi-layer SFM representations. Our final system achieves RMSE 22.36 on the development set and 24.98 on the evaluation set, ranking 1st on CPC3. These findings provide practical guidance for constructing SFM-based intrusive intelligibility predictors.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reference-aware SFM layers for intrusive intelligibility prediction
Yu, Hanlin
Zhou, Haoshuai
Cao, Boxuan
Mo, Changgeng
Li, Linkai
Wang, Shan X.
Audio and Speech Processing
Sound
Intrusive speech-intelligibility predictors that exploit explicit reference signals are now widespread, yet they have not consistently surpassed non-intrusive systems. We argue that a primary cause is the limited exploitation of speech foundation models (SFMs). This work revisits intrusive prediction by combining reference conditioning with multi-layer SFM representations. Our final system achieves RMSE 22.36 on the development set and 24.98 on the evaluation set, ranking 1st on CPC3. These findings provide practical guidance for constructing SFM-based intrusive intelligibility predictors.
title Reference-aware SFM layers for intrusive intelligibility prediction
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2509.17270