Ambisonics Super-Resolution Using A Waveform-Domain Neural Network
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866908475132477440 |
|---|---|
| author | Nawfal, Ismael Manias, Symeon Delikaris Souden, Mehrez Merimaa, Juha Atkins, Joshua McMullin, Elisabeth Pirhosseinloo, Shadi Phillips, Daniel |
| author_facet | Nawfal, Ismael Manias, Symeon Delikaris Souden, Mehrez Merimaa, Juha Atkins, Joshua McMullin, Elisabeth Pirhosseinloo, Shadi Phillips, Daniel |
| contents | Ambisonics is a spatial audio format describing a sound field. First-order Ambisonics (FOA) is a popular format comprising only four channels. This limited channel count comes at the expense of spatial accuracy. Ideally one would be able to take the efficiency of a FOA format without its limitations. We have devised a data-driven spatial audio solution that retains the efficiency of the FOA format but achieves quality that surpasses conventional renderers. Utilizing a fully convolutional time-domain audio neural network (Conv-TasNet), we created a solution that takes a FOA input and provides a higher order Ambisonics (HOA) output. This data driven approach is novel when compared to typical physics and psychoacoustic based renderers. Quantitative evaluations showed a 0.6dB average positional mean squared error difference between predicted and actual 3rd order HOA. The median qualitative rating showed an 80% improvement in perceived quality over the traditional rendering approach. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_00240 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Ambisonics Super-Resolution Using A Waveform-Domain Neural Network Nawfal, Ismael Manias, Symeon Delikaris Souden, Mehrez Merimaa, Juha Atkins, Joshua McMullin, Elisabeth Pirhosseinloo, Shadi Phillips, Daniel Audio and Speech Processing Sound Ambisonics is a spatial audio format describing a sound field. First-order Ambisonics (FOA) is a popular format comprising only four channels. This limited channel count comes at the expense of spatial accuracy. Ideally one would be able to take the efficiency of a FOA format without its limitations. We have devised a data-driven spatial audio solution that retains the efficiency of the FOA format but achieves quality that surpasses conventional renderers. Utilizing a fully convolutional time-domain audio neural network (Conv-TasNet), we created a solution that takes a FOA input and provides a higher order Ambisonics (HOA) output. This data driven approach is novel when compared to typical physics and psychoacoustic based renderers. Quantitative evaluations showed a 0.6dB average positional mean squared error difference between predicted and actual 3rd order HOA. The median qualitative rating showed an 80% improvement in perceived quality over the traditional rendering approach. |
| title | Ambisonics Super-Resolution Using A Waveform-Domain Neural Network |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2508.00240 |