Saved in:
Bibliographic Details
Main Authors: Götz, Philipp, Tuna, Cagdas, Brendel, Andreas, Walther, Andreas, Habets, Emanuël A. P.
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.19989
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • We present a method for blind acoustic parameter estimation from single-channel reverberant speech. The method is structured into three stages. In the first stage, a variational auto-encoder is trained to extract latent representations of acoustic impulse responses represented as mel-spectrograms. In the second stage, a separate speech encoder is trained to estimate low-dimensional representations from short segments of reverberant speech. Finally, the pre-trained speech encoder is combined with a small regression model and evaluated on two parameter regression tasks. Experimentally, the proposed method is shown to outperform a fully end-to-end trained baseline model.