Data Augmentation Using Neural Acoustic Fields With Retrieval-Augmented Pre-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ick, Christopher, Wichern, Gordon, Masuyama, Yoshiki, Germain, François G., Roux, Jonathan Le
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909585915248640
author Ick, Christopher
Wichern, Gordon
Masuyama, Yoshiki
Germain, François G.
Roux, Jonathan Le
author_facet Ick, Christopher
Wichern, Gordon
Masuyama, Yoshiki
Germain, François G.
Roux, Jonathan Le
contents This report details MERL's system for room impulse response (RIR) estimation submitted to the Generative Data Augmentation Workshop at ICASSP 2025 for Augmenting RIR Data (Task 1) and Improving Speaker Distance Estimation (Task 2). We first pre-train a neural acoustic field conditioned by room geometry on an external large-scale dataset in which pairs of RIRs and the geometries are provided. The neural acoustic field is then adapted to each target room by using the enrollment data, where we leverage either the provided room geometries or geometries retrieved from the external dataset, depending on availability. Lastly, we predict the RIRs for each pair of source and receiver locations specified by Task 1, and use these RIRs to train the speaker distance estimation model in Task 2.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14409
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data Augmentation Using Neural Acoustic Fields With Retrieval-Augmented Pre-training
Ick, Christopher
Wichern, Gordon
Masuyama, Yoshiki
Germain, François G.
Roux, Jonathan Le
Audio and Speech Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Sound
This report details MERL's system for room impulse response (RIR) estimation submitted to the Generative Data Augmentation Workshop at ICASSP 2025 for Augmenting RIR Data (Task 1) and Improving Speaker Distance Estimation (Task 2). We first pre-train a neural acoustic field conditioned by room geometry on an external large-scale dataset in which pairs of RIRs and the geometries are provided. The neural acoustic field is then adapted to each target room by using the enrollment data, where we leverage either the provided room geometries or geometries retrieved from the external dataset, depending on availability. Lastly, we predict the RIRs for each pair of source and receiver locations specified by Task 1, and use these RIRs to train the speaker distance estimation model in Task 2.
title Data Augmentation Using Neural Acoustic Fields With Retrieval-Augmented Pre-training
topic Audio and Speech Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Sound
url https://arxiv.org/abs/2504.14409