Ambisonics Binaural Rendering via Masked Magnitude Least Squares

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Berebi, Or, Brinkmann, Fabian, Weinzierl, Stefan, Rafaely, Boaz
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912450135195648
author Berebi, Or
Brinkmann, Fabian
Weinzierl, Stefan
Rafaely, Boaz
author_facet Berebi, Or
Brinkmann, Fabian
Weinzierl, Stefan
Rafaely, Boaz
contents Ambisonics rendering has become an integral part of 3D audio for headphones. It works well with existing recording hardware, the processing cost is mostly independent of the number of sound sources, and it elegantly allows for rotating the scene and listener. One challenge in Ambisonics headphone rendering is to find a perceptually well behaved low-order representation of the Head-Related Transfer Functions (HRTFs) that are contained in the rendering pipe-line. Low-order rendering is of interest, when working with microphone arrays containing only a few sensors, or for reducing the bandwidth for signal transmission. Magnitude Least Squares rendering became the de facto standard for this, which discards high-frequency interaural phase information in favor of reducing magnitude errors. Building upon this idea, we suggest Masked Magnitude Least Squares, which optimized the Ambisonics coefficients with a neural network and employs a spatio-spectral weighting mask to control the accuracy of the magnitude reconstruction. In the tested case, the weighting mask helped to maintain high-frequency notches in the low-order HRTFs and improved the modeled median plane localization performance in comparison to MagLS, while only marginally affecting the overall accuracy of the magnitude reconstruction.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18224
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ambisonics Binaural Rendering via Masked Magnitude Least Squares
Berebi, Or
Brinkmann, Fabian
Weinzierl, Stefan
Rafaely, Boaz
Audio and Speech Processing
Sound
Ambisonics rendering has become an integral part of 3D audio for headphones. It works well with existing recording hardware, the processing cost is mostly independent of the number of sound sources, and it elegantly allows for rotating the scene and listener. One challenge in Ambisonics headphone rendering is to find a perceptually well behaved low-order representation of the Head-Related Transfer Functions (HRTFs) that are contained in the rendering pipe-line. Low-order rendering is of interest, when working with microphone arrays containing only a few sensors, or for reducing the bandwidth for signal transmission. Magnitude Least Squares rendering became the de facto standard for this, which discards high-frequency interaural phase information in favor of reducing magnitude errors. Building upon this idea, we suggest Masked Magnitude Least Squares, which optimized the Ambisonics coefficients with a neural network and employs a spatio-spectral weighting mask to control the accuracy of the magnitude reconstruction. In the tested case, the weighting mask helped to maintain high-frequency notches in the low-order HRTFs and improved the modeled median plane localization performance in comparison to MagLS, while only marginally affecting the overall accuracy of the magnitude reconstruction.
title Ambisonics Binaural Rendering via Masked Magnitude Least Squares
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2501.18224