Ambisonics Super-Resolution Using A Waveform-Domain Neural Network

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nawfal, Ismael, Manias, Symeon Delikaris, Souden, Mehrez, Merimaa, Juha, Atkins, Joshua, McMullin, Elisabeth, Pirhosseinloo, Shadi, Phillips, Daniel
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908475132477440
author Nawfal, Ismael
Manias, Symeon Delikaris
Souden, Mehrez
Merimaa, Juha
Atkins, Joshua
McMullin, Elisabeth
Pirhosseinloo, Shadi
Phillips, Daniel
author_facet Nawfal, Ismael
Manias, Symeon Delikaris
Souden, Mehrez
Merimaa, Juha
Atkins, Joshua
McMullin, Elisabeth
Pirhosseinloo, Shadi
Phillips, Daniel
contents Ambisonics is a spatial audio format describing a sound field. First-order Ambisonics (FOA) is a popular format comprising only four channels. This limited channel count comes at the expense of spatial accuracy. Ideally one would be able to take the efficiency of a FOA format without its limitations. We have devised a data-driven spatial audio solution that retains the efficiency of the FOA format but achieves quality that surpasses conventional renderers. Utilizing a fully convolutional time-domain audio neural network (Conv-TasNet), we created a solution that takes a FOA input and provides a higher order Ambisonics (HOA) output. This data driven approach is novel when compared to typical physics and psychoacoustic based renderers. Quantitative evaluations showed a 0.6dB average positional mean squared error difference between predicted and actual 3rd order HOA. The median qualitative rating showed an 80% improvement in perceived quality over the traditional rendering approach.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00240
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ambisonics Super-Resolution Using A Waveform-Domain Neural Network
Nawfal, Ismael
Manias, Symeon Delikaris
Souden, Mehrez
Merimaa, Juha
Atkins, Joshua
McMullin, Elisabeth
Pirhosseinloo, Shadi
Phillips, Daniel
Audio and Speech Processing
Sound
Ambisonics is a spatial audio format describing a sound field. First-order Ambisonics (FOA) is a popular format comprising only four channels. This limited channel count comes at the expense of spatial accuracy. Ideally one would be able to take the efficiency of a FOA format without its limitations. We have devised a data-driven spatial audio solution that retains the efficiency of the FOA format but achieves quality that surpasses conventional renderers. Utilizing a fully convolutional time-domain audio neural network (Conv-TasNet), we created a solution that takes a FOA input and provides a higher order Ambisonics (HOA) output. This data driven approach is novel when compared to typical physics and psychoacoustic based renderers. Quantitative evaluations showed a 0.6dB average positional mean squared error difference between predicted and actual 3rd order HOA. The median qualitative rating showed an 80% improvement in perceived quality over the traditional rendering approach.
title Ambisonics Super-Resolution Using A Waveform-Domain Neural Network
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2508.00240