Generating Moving 3D Soundscapes with Latent Diffusion Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Templin, Christian, Zhu, Yanda, Wang, Hao
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914047339790336
author Templin, Christian
Zhu, Yanda
Wang, Hao
author_facet Templin, Christian
Zhu, Yanda
Wang, Hao
contents Spatial audio has become central to immersive applications such as VR/AR, cinema, and music. Existing generative audio models are largely limited to mono or stereo formats and cannot capture the full 3D localization cues available in first-order Ambisonics (FOA). Recent FOA models extend text-to-audio generation but remain restricted to static sources. In this work, we introduce SonicMotion, the first end-to-end latent diffusion framework capable of generating FOA audio with explicit control over moving sound sources. SonicMotion is implemented in two variations: 1) a descriptive model conditioned on natural language prompts, and 2) a parametric model conditioned on both text and spatial trajectory parameters for higher precision. To support training and evaluation, we construct a new dataset of over one million simulated FOA caption pairs that include both static and dynamic sources with annotated azimuth, elevation, and motion attributes. Experiments show that SonicMotion achieves state-of-the-art semantic alignment and perceptual quality comparable to leading text-to-audio systems, while uniquely attaining low spatial localization error.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07318
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generating Moving 3D Soundscapes with Latent Diffusion Models
Templin, Christian
Zhu, Yanda
Wang, Hao
Sound
Artificial Intelligence
Audio and Speech Processing
Spatial audio has become central to immersive applications such as VR/AR, cinema, and music. Existing generative audio models are largely limited to mono or stereo formats and cannot capture the full 3D localization cues available in first-order Ambisonics (FOA). Recent FOA models extend text-to-audio generation but remain restricted to static sources. In this work, we introduce SonicMotion, the first end-to-end latent diffusion framework capable of generating FOA audio with explicit control over moving sound sources. SonicMotion is implemented in two variations: 1) a descriptive model conditioned on natural language prompts, and 2) a parametric model conditioned on both text and spatial trajectory parameters for higher precision. To support training and evaluation, we construct a new dataset of over one million simulated FOA caption pairs that include both static and dynamic sources with annotated azimuth, elevation, and motion attributes. Experiments show that SonicMotion achieves state-of-the-art semantic alignment and perceptual quality comparable to leading text-to-audio systems, while uniquely attaining low spatial localization error.
title Generating Moving 3D Soundscapes with Latent Diffusion Models
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2507.07318