Should Audio Front-ends be Adaptive? Comparing Learnable and Adaptive Front-ends
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917913998393344 |
|---|---|
| author | Zhang, Qiquan Wickramasinghe, Buddhi Ambikairajah, Eliathamby Sethu, Vidhyasaharan Li, Haizhou |
| author_facet | Zhang, Qiquan Wickramasinghe, Buddhi Ambikairajah, Eliathamby Sethu, Vidhyasaharan Li, Haizhou |
| contents | Hand-crafted features, such as Mel-filterbanks, have traditionally been the choice for many audio processing applications. Recently, there has been a growing interest in learnable front-ends that extract representations directly from the raw audio waveform. \textcolor{black}{However, both hand-crafted filterbanks and current learnable front-ends lead to fixed computation graphs at inference time, failing to dynamically adapt to varying acoustic environments, a key feature of human auditory systems.} To this end, we explore the question of whether audio front-ends should be adaptive by comparing the Ada-FE front-end (a recently developed adaptive front-end that employs a neural adaptive feedback controller to dynamically adjust the Q-factors of its spectral decomposition filters) to established learnable front-ends. Specifically, we systematically investigate learnable front-ends and Ada-FE across two commonly used back-end backbones and a wide range of audio benchmarks including speech, sound event, and music. The comprehensive results show that our Ada-FE outperforms advanced learnable front-ends, and more importantly, it exhibits impressive stability or robustness on test samples over various training epochs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_03260 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Should Audio Front-ends be Adaptive? Comparing Learnable and Adaptive Front-ends Zhang, Qiquan Wickramasinghe, Buddhi Ambikairajah, Eliathamby Sethu, Vidhyasaharan Li, Haizhou Audio and Speech Processing Sound Hand-crafted features, such as Mel-filterbanks, have traditionally been the choice for many audio processing applications. Recently, there has been a growing interest in learnable front-ends that extract representations directly from the raw audio waveform. \textcolor{black}{However, both hand-crafted filterbanks and current learnable front-ends lead to fixed computation graphs at inference time, failing to dynamically adapt to varying acoustic environments, a key feature of human auditory systems.} To this end, we explore the question of whether audio front-ends should be adaptive by comparing the Ada-FE front-end (a recently developed adaptive front-end that employs a neural adaptive feedback controller to dynamically adjust the Q-factors of its spectral decomposition filters) to established learnable front-ends. Specifically, we systematically investigate learnable front-ends and Ada-FE across two commonly used back-end backbones and a wide range of audio benchmarks including speech, sound event, and music. The comprehensive results show that our Ada-FE outperforms advanced learnable front-ends, and more importantly, it exhibits impressive stability or robustness on test samples over various training epochs. |
| title | Should Audio Front-ends be Adaptive? Comparing Learnable and Adaptive Front-ends |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2502.03260 |