Deep Learning for Personalized Binaural Audio Reproduction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lu, Xikun, Chen, Yunda, Chen, Zehua, Wang, Jie, Liu, Mingxing, Hu, Hongmei, Zheng, Chengshi, Bleeck, Stefan, Sang, Jinqiu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909761193115648
author Lu, Xikun
Chen, Yunda
Chen, Zehua
Wang, Jie
Liu, Mingxing
Hu, Hongmei
Zheng, Chengshi
Bleeck, Stefan
Sang, Jinqiu
author_facet Lu, Xikun
Chen, Yunda
Chen, Zehua
Wang, Jie
Liu, Mingxing
Hu, Hongmei
Zheng, Chengshi
Bleeck, Stefan
Sang, Jinqiu
contents Personalized binaural audio reproduction is the basis of realistic spatial localization, sound externalization, and immersive listening, directly shaping user experience and listening effort. This survey reviews recent advances in deep learning for this task and organizes them by generation mechanism into two paradigms: explicit personalized filtering and end-to-end rendering. Explicit methods predict personalized head-related transfer functions (HRTFs) from sparse measurements, morphological features, or environmental cues, and then use them in the conventional rendering pipeline. End-to-end methods map source signals directly to binaural signals, aided by other inputs such as visual, textual, or parametric guidance, and they learn personalization within the model. We also summarize the field's main datasets and evaluation metrics to support fair and repeatable comparison. Finally, we conclude with a discussion of key applications enabled by these technologies, current technical limitations, and potential research directions for deep learning-based spatial audio systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00400
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deep Learning for Personalized Binaural Audio Reproduction
Lu, Xikun
Chen, Yunda
Chen, Zehua
Wang, Jie
Liu, Mingxing
Hu, Hongmei
Zheng, Chengshi
Bleeck, Stefan
Sang, Jinqiu
Audio and Speech Processing
Sound
Personalized binaural audio reproduction is the basis of realistic spatial localization, sound externalization, and immersive listening, directly shaping user experience and listening effort. This survey reviews recent advances in deep learning for this task and organizes them by generation mechanism into two paradigms: explicit personalized filtering and end-to-end rendering. Explicit methods predict personalized head-related transfer functions (HRTFs) from sparse measurements, morphological features, or environmental cues, and then use them in the conventional rendering pipeline. End-to-end methods map source signals directly to binaural signals, aided by other inputs such as visual, textual, or parametric guidance, and they learn personalization within the model. We also summarize the field's main datasets and evaluation metrics to support fair and repeatable comparison. Finally, we conclude with a discussion of key applications enabled by these technologies, current technical limitations, and potential research directions for deep learning-based spatial audio systems.
title Deep Learning for Personalized Binaural Audio Reproduction
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2509.00400