PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Vosoughi, Ali, Zang, Yongyi, Yang, Qihui, Paek, Nathan, Leistikow, Randal, Xu, Chenliang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914120903688192
author Vosoughi, Ali
Zang, Yongyi
Yang, Qihui
Paek, Nathan
Leistikow, Randal
Xu, Chenliang
author_facet Vosoughi, Ali
Zang, Yongyi
Yang, Qihui
Paek, Nathan
Leistikow, Randal
Xu, Chenliang
contents Room impulse response (RIR) generation remains a critical challenge for creating immersive virtual acoustic environments. Current methods suffer from two fundamental limitations: the scarcity of full-band RIR datasets and the inability of existing models to generate acoustically accurate responses from diverse input modalities. We present PromptReverb, a two-stage generative framework that addresses these challenges. Our approach combines a variational autoencoder that upsamples band-limited RIRs to full-band quality (48 kHz), and a conditional diffusion transformer model based on rectified flow matching that generates RIRs from descriptions in natural language. Empirical evaluation demonstrates that PromptReverb produces RIRs with superior perceptual quality and acoustic accuracy compared to existing methods, achieving 8.8% mean RT60 error compared to -37% for widely used baselines and yielding more realistic room-acoustic parameters. Our method enables practical applications in virtual reality, architectural acoustics, and audio production where flexible, high-quality RIR synthesis is essential.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
Vosoughi, Ali
Zang, Yongyi
Yang, Qihui
Paek, Nathan
Leistikow, Randal
Xu, Chenliang
Sound
Artificial Intelligence
I.2.6, H.5.5
Room impulse response (RIR) generation remains a critical challenge for creating immersive virtual acoustic environments. Current methods suffer from two fundamental limitations: the scarcity of full-band RIR datasets and the inability of existing models to generate acoustically accurate responses from diverse input modalities. We present PromptReverb, a two-stage generative framework that addresses these challenges. Our approach combines a variational autoencoder that upsamples band-limited RIRs to full-band quality (48 kHz), and a conditional diffusion transformer model based on rectified flow matching that generates RIRs from descriptions in natural language. Empirical evaluation demonstrates that PromptReverb produces RIRs with superior perceptual quality and acoustic accuracy compared to existing methods, achieving 8.8% mean RT60 error compared to -37% for widely used baselines and yielding more realistic room-acoustic parameters. Our method enables practical applications in virtual reality, architectural acoustics, and audio production where flexible, high-quality RIR synthesis is essential.
title PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
topic Sound
Artificial Intelligence
I.2.6, H.5.5
url https://arxiv.org/abs/2510.22439