Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chung, Soo-Whan, Choi, Min-Seok
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911102906925056
author Chung, Soo-Whan
Choi, Min-Seok
author_facet Chung, Soo-Whan
Choi, Min-Seok
contents This paper introduces a novel approach to speech restoration by integrating a context-related conditioning strategy. Specifically, we employ the diffusion-based generative restoration model, UNIVERSE++, as a backbone to evaluate the effectiveness of contextual representations. We incorporate acoustic context embeddings extracted from the CLAP model, which capture the environmental attributes of input audio. Additionally, we propose an Acoustic Context (ACX) representation that refines CLAP embeddings to better handle various distortion factors and their intensity in speech signals. Unlike content-based approaches that rely on linguistic and speaker attributes, ACX provides contextual information that enables the restoration model to distinguish and mitigate distortions better. Experimental results indicate that context-aware conditioning improves both restoration performance and its stability across diverse distortion conditions, reducing variability compared to content-based methods.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08953
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
Chung, Soo-Whan
Choi, Min-Seok
Audio and Speech Processing
Sound
This paper introduces a novel approach to speech restoration by integrating a context-related conditioning strategy. Specifically, we employ the diffusion-based generative restoration model, UNIVERSE++, as a backbone to evaluate the effectiveness of contextual representations. We incorporate acoustic context embeddings extracted from the CLAP model, which capture the environmental attributes of input audio. Additionally, we propose an Acoustic Context (ACX) representation that refines CLAP embeddings to better handle various distortion factors and their intensity in speech signals. Unlike content-based approaches that rely on linguistic and speaker attributes, ACX provides contextual information that enables the restoration model to distinguish and mitigate distortions better. Experimental results indicate that context-aware conditioning improves both restoration performance and its stability across diverse distortion conditions, reducing variability compared to content-based methods.
title Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2508.08953