No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Haoshuai, Mo, Changgeng, Cao, Boxuan, Li, Linkai, Wang, Shan Xiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910982256721920
author Zhou, Haoshuai
Mo, Changgeng
Cao, Boxuan
Li, Linkai
Wang, Shan Xiang
author_facet Zhou, Haoshuai
Mo, Changgeng
Cao, Boxuan
Li, Linkai
Wang, Shan Xiang
contents Personalized speech intelligibility prediction is challenging. Previous approaches have mainly relied on audiograms, which are inherently limited in accuracy as they only capture a listener's hearing threshold for pure tones. Rather than incorporating additional listener features, we propose a novel approach that leverages an individual's existing intelligibility data to predict their performance on new audio. We introduce the Support Sample-Based Intelligibility Prediction Network (SSIPNet), a deep learning model that leverages speech foundation models to build a high-dimensional representation of a listener's speech recognition ability from multiple support (audio, score) pairs, enabling accurate predictions for unseen audio. Results on the Clarity Prediction Challenge dataset show that, even with a small number of support (audio, score) pairs, our method outperforms audiogram-based predictions. Our work presents a new paradigm for personalized speech intelligibility prediction.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02039
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction
Zhou, Haoshuai
Mo, Changgeng
Cao, Boxuan
Li, Linkai
Wang, Shan Xiang
Audio and Speech Processing
Artificial Intelligence
Sound
Personalized speech intelligibility prediction is challenging. Previous approaches have mainly relied on audiograms, which are inherently limited in accuracy as they only capture a listener's hearing threshold for pure tones. Rather than incorporating additional listener features, we propose a novel approach that leverages an individual's existing intelligibility data to predict their performance on new audio. We introduce the Support Sample-Based Intelligibility Prediction Network (SSIPNet), a deep learning model that leverages speech foundation models to build a high-dimensional representation of a listener's speech recognition ability from multiple support (audio, score) pairs, enabling accurate predictions for unseen audio. Results on the Clarity Prediction Challenge dataset show that, even with a small number of support (audio, score) pairs, our method outperforms audiogram-based predictions. Our work presents a new paradigm for personalized speech intelligibility prediction.
title No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2506.02039