Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Masuyama, Yoshiki, Germain, Francois G., Wichern, Gordon, Hori, Chiori, Roux, Jonathan Le
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917359224094720
author Masuyama, Yoshiki
Germain, Francois G.
Wichern, Gordon
Hori, Chiori
Roux, Jonathan Le
author_facet Masuyama, Yoshiki
Germain, Francois G.
Wichern, Gordon
Hori, Chiori
Roux, Jonathan Le
contents First-order Ambisonics (FOA) is a standard spatial audio format based on spherical harmonic decomposition. Its zeroth- and first-order components capture the sound pressure and particle velocity, respectively. Recently, physics-informed neural networks have been applied to the spatial interpolation of FOA signals, regularizing the network outputs based on soft penalty terms derived from physical principles, e.g., the linearized momentum equation. In this paper, we reformulate the task so that the predicted FOA signal automatically satisfies the linearized momentum equation. Our network approximates a scalar function called velocity potential, rather than the FOA signal itself. Then, the FOA signal can be readily recovered through the partial derivatives of the velocity potential with respect to the network inputs (i.e., time and microphone position) according to physics of sound propagation. By deriving the four channels of FOA from the single-channel velocity potential, the reconstructed signal follows the physical principle at any time and position by construction. Experimental results on room impulse response reconstruction confirm the effectiveness of the proposed framework.
format Preprint
id arxiv_https___arxiv_org_abs_2603_22589
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
Masuyama, Yoshiki
Germain, Francois G.
Wichern, Gordon
Hori, Chiori
Roux, Jonathan Le
Sound
Audio and Speech Processing
Signal Processing
First-order Ambisonics (FOA) is a standard spatial audio format based on spherical harmonic decomposition. Its zeroth- and first-order components capture the sound pressure and particle velocity, respectively. Recently, physics-informed neural networks have been applied to the spatial interpolation of FOA signals, regularizing the network outputs based on soft penalty terms derived from physical principles, e.g., the linearized momentum equation. In this paper, we reformulate the task so that the predicted FOA signal automatically satisfies the linearized momentum equation. Our network approximates a scalar function called velocity potential, rather than the FOA signal itself. Then, the FOA signal can be readily recovered through the partial derivatives of the velocity potential with respect to the network inputs (i.e., time and microphone position) according to physics of sound propagation. By deriving the four channels of FOA from the single-channel velocity potential, the reconstructed signal follows the physical principle at any time and position by construction. Experimental results on room impulse response reconstruction confirm the effectiveness of the proposed framework.
title Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
topic Sound
Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2603.22589