On HRTF Notch Frequency Prediction Using Anthropometric Features and Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arbel, Lior, Ananthabhotla, Ishwarya, Ben-Hur, Zamir, Alon, David Lou, Rafaely, Boaz
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910913711308800
author Arbel, Lior
Ananthabhotla, Ishwarya
Ben-Hur, Zamir
Alon, David Lou
Rafaely, Boaz
author_facet Arbel, Lior
Ananthabhotla, Ishwarya
Ben-Hur, Zamir
Alon, David Lou
Rafaely, Boaz
contents High fidelity spatial audio often performs better when produced using a personalized head-related transfer function (HRTF). However, the direct acquisition of HRTFs is cumbersome and requires specialized equipment. Thus, many personalization methods estimate HRTF features from easily obtained anthropometric features of the pinna, head, and torso. The first HRTF notch frequency (N1) is known to be a dominant feature in elevation localization, and thus a useful feature for HRTF personalization. This paper describes the prediction of N1 frequency from pinna anthropometry using a neural model. Prediction is performed separately on three databases, both simulated and measured, and then by domain mixing in-between the databases. The model successfully predicts N1 frequency for individual databases and by domain mixing between some databases. Prediction errors are better or comparable to those previously reported, showing significant improvement when acquired over a large database and with a larger output range.
format Preprint
id arxiv_https___arxiv_org_abs_2403_07579
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On HRTF Notch Frequency Prediction Using Anthropometric Features and Neural Networks
Arbel, Lior
Ananthabhotla, Ishwarya
Ben-Hur, Zamir
Alon, David Lou
Rafaely, Boaz
Audio and Speech Processing
High fidelity spatial audio often performs better when produced using a personalized head-related transfer function (HRTF). However, the direct acquisition of HRTFs is cumbersome and requires specialized equipment. Thus, many personalization methods estimate HRTF features from easily obtained anthropometric features of the pinna, head, and torso. The first HRTF notch frequency (N1) is known to be a dominant feature in elevation localization, and thus a useful feature for HRTF personalization. This paper describes the prediction of N1 frequency from pinna anthropometry using a neural model. Prediction is performed separately on three databases, both simulated and measured, and then by domain mixing in-between the databases. The model successfully predicts N1 frequency for individual databases and by domain mixing between some databases. Prediction errors are better or comparable to those previously reported, showing significant improvement when acquired over a large database and with a larger output range.
title On HRTF Notch Frequency Prediction Using Anthropometric Features and Neural Networks
topic Audio and Speech Processing
url https://arxiv.org/abs/2403.07579