Matching Reverberant Speech Through Learned Acoustic Embeddings and Feedback Delay Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Götz, Philipp, Santo, Gloria Dal, Schlecht, Sebastian J., Välimäki, Vesa, Habets, Emanuël A. P.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911234551447552
author Götz, Philipp
Santo, Gloria Dal
Schlecht, Sebastian J.
Välimäki, Vesa
Habets, Emanuël A. P.
author_facet Götz, Philipp
Santo, Gloria Dal
Schlecht, Sebastian J.
Välimäki, Vesa
Habets, Emanuël A. P.
contents Reverberation conveys critical acoustic cues about the environment, supporting spatial awareness and immersion. For auditory augmented reality (AAR) systems, generating perceptually plausible reverberation in real time remains a key challenge, especially when explicit acoustic measurements are unavailable. We address this by formulating blind estimation of artificial reverberation parameters as a reverberant signal matching task, leveraging a learned room-acoustic prior. Furthermore, we propose a feedback delay network (FDN) structure that reproduces both frequency-dependent decay times and the direct-to-reverberation ratio of a target space. Experimental evaluation against a leading automatic FDN tuning method demonstrates improvements in estimated room-acoustic parameters and perceptual plausibility of artificial reverberant speech. These results highlight the potential of our approach for efficient, perceptually consistent reverberation rendering in AAR applications.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23158
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Matching Reverberant Speech Through Learned Acoustic Embeddings and Feedback Delay Networks
Götz, Philipp
Santo, Gloria Dal
Schlecht, Sebastian J.
Välimäki, Vesa
Habets, Emanuël A. P.
Audio and Speech Processing
Reverberation conveys critical acoustic cues about the environment, supporting spatial awareness and immersion. For auditory augmented reality (AAR) systems, generating perceptually plausible reverberation in real time remains a key challenge, especially when explicit acoustic measurements are unavailable. We address this by formulating blind estimation of artificial reverberation parameters as a reverberant signal matching task, leveraging a learned room-acoustic prior. Furthermore, we propose a feedback delay network (FDN) structure that reproduces both frequency-dependent decay times and the direct-to-reverberation ratio of a target space. Experimental evaluation against a leading automatic FDN tuning method demonstrates improvements in estimated room-acoustic parameters and perceptual plausibility of artificial reverberant speech. These results highlight the potential of our approach for efficient, perceptually consistent reverberation rendering in AAR applications.
title Matching Reverberant Speech Through Learned Acoustic Embeddings and Feedback Delay Networks
topic Audio and Speech Processing
url https://arxiv.org/abs/2510.23158